Pieter Abbeel
pieter abbeel teaches robots to learn from observation, bridging the gap between academic insight and real-world mechanical intelligence
Pieter Abbeel (born 1977) is a professor of electrical engineering and computer sciences, Director of the Berkeley Robot Learning Lab, and co-director of the Berkeley AI Research (BAIR) Lab at the University of California, Berkeley. He is also the co-founder of Covariant, a venture-funded start-up that aims to teach… wikipedia →
12-month trajectory
interviews & talks

Pieter Abbeel: Deep Reinforcement Learning | Lex Fridman Podcast #10

Pieter Abbeel explains our approach to building our #generalizedAI at @fortune’s Brainstorm #AI

Pieter Abbeel Interview

Pieter Abbeel - Really Quick Questions with a Berkeley Professor

S3 E9 Geoff Hinton, the "Godfather of AI", quits Google to warn of AI risks (Host: Pieter Abbeel)
recent news
amazon ai strategy shift
- Amazon Reorganises AI, Deprecates Models, Consolidates Under Abbeel - TechGig
- Amazon shutters AGI lab to focus on frontier model initiative - NewsBytes
- After layoffs in AGI division, shutdown of Lab, Amazon is reorganizing teams and refocusing its AI strate - The Times of India
+ 5 more
amazon leadership changes
- Amazon winds down most flagship AI models in strategy overhaul — report - The Edge Malaysia
- Amazon scales back flagship AI models in major strategy shift - The News International
- Amazon Kills Most Nova AI Models, Bets Big on a Single Flagship Before Earnings - finance.biggo.com
+ 3 more
amazon robotics initiatives
- Amazon Pivots From Nova Models To A New Frontier Bet - Finimize
- Amazon scales back flagship AI models in major strategy shift - The News International
- Amazon AI Strategy Reportedly Shifts Before Earnings, Nova AI Models Put In ‘Keep The Lights On’ Mode - Yahoo Finance
+ 2 more
abbeel ai risk research
- Who are the Execs Leading Amazon’s Future AI Strategy? - Business Chief
- Fortune Tech: Amazon’s new AI leader, Oracle’s stalled plans, Oscars on YouTube - Fortune
- Amazon AI chief Rohit Prasad resigns amid restructuring of AGI unit - The American Bazaar
+ 8 more
abbeel robotics research
- Amazon CEO Andy Jassy announces departure of AI exec Rohit Prasad in leadership shake-up - Fortune
- Amazon shakes up AI team as veteran Prasad leaves, DeSantis promoted - Reuters
- Industry Insights: How Amazon Leveraged Covariant's Technology to Accelerate Robot Development - A3 Association for Advancing Automation
+ 8 more
abbeel startup ventures
- An update on how we’re accelerating the use of AI in robotics at scale - About Amazon
- Managing extreme AI risks amid rapid progress - Science | AAAS
- This company is using the AI that powers ChatGPT to help warehouse robots expect the unexpected - Business Insider
+ 12 more
unrelated headlines
dispatch
Amazon Bets Everything on Abbeel's Vision
Pieter Abbeel is the Belgian-born Berkeley roboticist — Andrew Ng's first PhD student, co-founder of Covariant and Gradescope, and the teacher behind a startling number of today's AI founders — whom Amazon has just put in charge of its most important model bet. He's on the index this week because Amazon is winding down most of its Nova AI models into maintenance mode and pouring the talent and compute into a single frontier effort he leads. The takeaway for a CTO: one of the largest cloud vendors just admitted a broad model lineup doesn't win, chose depth over breadth, and quietly reminded everyone that vendor models can be deprecated out from under you.
Start with a helicopter flying upside down.
Not a stunt pilot. A computer, flying a real radio-controlled helicopter through maneuvers that only a handful of human experts in the world can pull off — barrel rolls, tail slides, a move pilots call "chaos" where the aircraft tumbles in a way that looks like it's about to fall out of the sky, and then recovers. The machine wasn't told how to do any of that with equations. It learned by watching a human do it, over and over, and working backward from what it saw to figure out what the pilot must have been trying to achieve.
The person behind that work is Pieter Abbeel. And that helicopter is the whole man in one image. Because his entire career is about one stubborn idea: that the way to make machines competent in the physical world is not to program them, but to let them learn — from data, from demonstration, from watching.
Abbeel was born in 1977, in Antwerp, in Belgium. He grew up in Brasschaat, a suburb just outside the city. He studied electrical engineering at KU Leuven, one of the oldest universities in Europe, and then he did what a lot of ambitious young engineers did around the turn of the millennium — he came to the United States, to Stanford, for a PhD.
At Stanford he became the first PhD student of a professor named Andrew Ng. If you follow this field, you know that name. Ng went on to co-found Google Brain and Coursera, and to become one of the most influential teachers in modern AI. Abbeel was his first. That helicopter work was the thesis. The formal title was apprenticeship learning and reinforcement learning with application to robotic control, and he finished it in 2008.
Now, sit with that word: apprenticeship. It tells you how he thinks. An apprentice doesn't get handed a rulebook. An apprentice stands next to a master and copies, and slowly absorbs the judgment behind the moves. Abbeel spent twenty years trying to give machines that ability.
In 2008 he joined the faculty at UC Berkeley, where he still is — director of the Robot Learning Lab, and co-director of the broader Berkeley AI Research group, which everyone just calls BAIR. And this is where the story gets bigger than any one person, because Abbeel turned out to be one of the great teachers of his generation. Not in the sense of lecturing. In the sense of producing people.
Look at who came through his lab. John Schulman, who co-founded OpenAI and was one of the lead architects of ChatGPT. Aravind Srinivas, who co-founded Perplexity. Chelsea Finn and Sergey Levine, who co-founded Physical Intelligence, one of the hottest robotics startups in the world right now. Deepak Pathak, who founded Skild. Jonathan Ho, who was central to the invention of diffusion models — the technique underneath basically every AI image generator you've used. Berkeley's own count puts it at more than a dozen AI companies founded by Abbeel and his students. If you wanted to draw the family tree of modern AI, a surprising number of the branches run straight through this one lab.
He didn't just stay in the academy, though. In 2014 he co-founded Gradescope, a tool that uses AI to help professors grade exams, and it ended up in hundreds of universities. Then in 2016 he spent a year as a research scientist at OpenAI, in the early days, working on robotics and on a problem we'll come back to — how you train a robot in simulation and then get it to work in the real world.
And in 2017 he started the company that matters most for today's story. It's called Covariant. He founded it with three other researchers out of that same OpenAI and Berkeley orbit — Peter Chen, Rocky Duan, and Tianhao Zhang. The pitch was deceptively simple. Warehouses are full of robot arms that are strong and precise and completely blind to anything they weren't explicitly programmed for. Covariant wanted to build the brain — one AI system that could let a robot arm look into a bin of random objects it had never seen and just pick the right one. Universal, learned, adaptable.
Covariant raised a couple hundred million dollars over its life, from serious investors, and it put real robots into real distribution centers. And in March of 2024, at a logistics trade show, the company unveiled something it called a robotics foundation model. We'll get into what that actually means later, because it's the key to everything. For now, just hold onto the phrase. A foundation model — for robots.
That's the man. Belgian engineer, Andrew Ng's first student, a teacher who seeded half of Silicon Valley's AI labs, a founder who spent seven years trying to give warehouse robots something like common sense. Which brings us to why he's in the news this week, and it's not for a robot at all.
Here's what happened.
In 2024, Amazon did a deal with Covariant. Not a straight acquisition — the structure was a licensing agreement for Covariant's AI models, plus hiring a big chunk of the team. Reporting puts it at roughly a quarter of Covariant's people. Abbeel came over with them. This has become a common move for large tech companies — you license the technology and absorb the talent without formally buying the company. Amazon has done versions of it more than once.
And then Abbeel started climbing inside Amazon, fast. At the end of 2024, Amazon stood up something called the AGI Lab in San Francisco. AGI — artificial general intelligence, the industry's term for AI that reaches human-level breadth. The lab was built largely around people Amazon had hired from another startup, a company called Adept, including Adept's co-founder and CEO, David Luan. Abbeel and Luan were named as the leaders on the project. According to The Information, the lab grew to around eighty people at its peak.
That was the ambition a year and a half ago. Here's where it is now.
David Luan left Amazon in February. More than a dozen of those Adept hires left with him, or around the same time. The executive who'd been running Amazon's whole AGI effort since 2023, Rohit Prasad, was reported to be leaving at the end of last year. Andy Jassy, Amazon's CEO, reorganized the entire AI group and put it under Peter DeSantis — a 27-year Amazon veteran whose background is in AWS infrastructure and custom chips, not frontier research. And this week, Amazon confirmed it's closing that San Francisco AGI site entirely, as part of layoffs across the AGI organization.
So a lab loses its founding leader. Then the broader boss leaves. Then the lab shuts. As one write-up put it, that's not a quiet reorg — that's a strategy changing shape in public.
But the closure is only half the news. The other half broke through reporting from Business Insider in late July, and it's the part that actually explains Abbeel's role.
Amazon has a family of homegrown AI models it calls Nova — text models, a multimodal model, an image generator, a video generator. Amazon is now winding most of them down. According to that reporting, four of the flagship Nova models — the top-end language model, the multimodal one, the video generator, and the image generator — are being moved into maintenance-only mode. Inside Amazon, employees reportedly describe this as "KTLO" — keep the lights on. The products keep running for existing customers, but active development stops. Nobody's making them better anymore.
And the reason they're being wound down is the thing Abbeel now runs. Amazon has created a new group called Frontier Model Research, and it has become the company's single highest AI priority. The engineers and the computing power that used to be spread across all those Nova models are being funneled into it. The goal is not a broad lineup. The goal is one model — one genuinely competitive frontier model — and Pieter Abbeel is leading the effort to build it. Reporting suggests a new flagship could be unveiled at Amazon's re:Invent conference later this year, and it may even keep the Nova name.
An Amazon spokesperson framed it to Reuters in the language you'd expect. They said the company is, quote, "sharpening our focus on the initiatives that matter most for customers, so we can move faster on what counts." That's the corporate translation. The plain-English translation is: the scattershot approach didn't work, and we're betting the concentration on one man's team.
So why him? Why hand the most important model bet at a trillion-dollar company to a robotics professor? To answer that, we have to go under the hood — into what Abbeel actually knows how to do, and why it might be exactly what a frontier model needs right now.
Let's start with the thing he's spent his life on: reinforcement learning.
Most of the AI you interact with is trained by imitation. You show a model billions of examples — text, images — and it learns to predict the next piece. That's how a large language model works at its core. It's an extraordinarily good mimic of patterns in data that already exists.
Reinforcement learning is different, and it's older in spirit. In reinforcement learning, you don't show the system the right answer. You put an agent in an environment, let it try things, and give it a reward signal when it does well. It learns by trial and error, the way you'd train a dog, or the way you learned to ride a bike — not from a manual, but from falling over and adjusting. That's the technique behind the systems that beat humans at Go and at video games. It's powerful, and it's notoriously finicky, because the agent has to explore an enormous space of possible actions and figure out which ones actually pay off, often with the reward arriving long after the action that earned it.
Now, Abbeel's specific twist on this is important. Reinforcement learning needs a reward function — a definition of what "good" means. For a lot of real tasks, nobody can write that definition down. How do you write an equation for "flew the helicopter beautifully"? You can't. So Abbeel worked on the inverse problem. Instead of specifying the reward and searching for the behavior, you observe the behavior — the expert human — and you work backward to infer the reward that the expert must have been optimizing. That's inverse reinforcement learning, and it's the technical heart of apprenticeship learning. Watch the master, deduce what they were trying to do, then learn to do it yourself. That upside-down helicopter was the proof.
Here's why this matters far beyond helicopters. The one-line summary a lot of people use for Abbeel is that he teaches robots to learn from observation — to bridge academic insight and real-world mechanical intelligence. And the reason that's hard, the reason it's the central problem in robotics, comes down to data.
Language models had the internet. There are trillions of words sitting on the web, for free, ready to be fed into a model. Robots have nothing like that. There is no internet-scale archive of a robot arm picking up ten billion different objects. Every piece of robot training data has to be generated — by a real robot, moving in real time, in the real world, which is slow and expensive and doesn't scale the way scraping a website does. This is the data problem, and it's the wall that robotics keeps running into.
Abbeel's career is essentially a series of attacks on that wall. One of his most influential ideas, from his OpenAI period, is called domain randomization. The insight goes like this. You can't easily collect real-world robot data, but you can generate infinite data in simulation, for almost nothing. The catch is that simulations are never quite right — the friction is a little off, the lighting is wrong, the physics is approximate — so a robot trained in simulation usually falls apart in reality. That gap even has a name in the field: the sim-to-real gap. Domain randomization is a clever trick to close it. Instead of trying to make one perfect simulation, you deliberately randomize thousands of imperfect ones — vary the colors, the weights, the friction, the lighting, all over the place. A model trained across that chaos stops relying on any single detail being correct. It learns the robust underlying skill. And then, when it meets the real world, the real world just looks like one more variation it's already seen. That idea helped make it possible to train a robot hand in simulation and have it work on real hardware.
Now put those two threads together — learning from limited data, and learning to act rather than just to predict — and you can see why Amazon's bet is less strange than it first looks.
Because the frontier of AI has moved. For a few years, progress was mostly about scale — bigger models, more text, predict the next word better. That era is hitting diminishing returns, partly because the industry is running out of fresh internet to train on. The new frontier is about models that reason, that take actions, that use tools, that improve themselves without a human labeling every example. In other words, the new frontier is starting to look a lot like reinforcement learning and agents — which is precisely Abbeel's home turf. The techniques that made the latest reasoning models good are, at their core, reinforcement learning techniques. And one of the people who wrote the foundational papers on modern reinforcement learning, John Schulman, came out of Abbeel's lab.
There's one more piece, and it's the robotics foundation model we set aside earlier — the thing Covariant unveiled in 2024. The idea there was to take the recipe that made language models work and point it at robots. Build one large model trained not just on text, but on text and images and video and the actual physical actions and sensor readings of robots doing tasks. Ask it, in plain language, to do something, and have it produce the right physical motion. It's an attempt to give a robot the kind of broad, transferable understanding that ChatGPT has for language — a general prior about how the physical world behaves, so the robot isn't learning every new task from scratch. Whether that fully works at scale is still an open question. But it's the same instinct as everything else in his career: find the general learning method, feed it the right data, and don't hand-program the specifics.
So the person Amazon put in charge isn't a language-model specialist who wandered in from robotics. He's someone whose entire toolkit — reinforcement learning, learning from demonstration, squeezing capability out of scarce data, transferring skills across domains — happens to be the toolkit the frontier is now demanding. That doesn't guarantee the bet pays off. But it's a coherent bet, not a panicked one.
Which brings us to the part that matters if you're building a company, or running an engineering org, and watching all this from the outside. What do you actually take from it?
Start with the biggest signal, because Amazon just said something out loud that a lot of people suspected. A wide lineup of models does not win. For two years the reflex across the industry was breadth — ship a model for every task, a text one, an image one, a video one, put them all in the catalog. Amazon built exactly that with Nova, and Amazon is now walking away from most of it to concentrate everything on a single frontier model. When one of the three largest cloud providers on earth decides that depth beats breadth, that's not a rumor. That's a data point about where the economics actually are. Building and maintaining a frontier model is so expensive, and the gap between the best model and the second-best is so consequential, that spreading yourself thin is a way to lose slowly.
Second, and this one should make every CTO sit up: those Nova models didn't get killed. They got moved to "keep the lights on." Think about what that means if you're a customer who built a product on one of them. The model still runs. Your integration still works. But nobody is improving it, nobody is patching its weaknesses, and the roadmap is quietly gone. That is the specific risk of building on someone else's model, and Amazon just gave you a live example of how it plays out. It rarely arrives as a dramatic shutdown notice. It arrives as maintenance mode. So when you pick a vendor model to build on, the question isn't only "how good is it today." The question is "how central is this to the vendor's actual strategy," because the peripheral models are the ones that go quiet first. Design for portability. Keep the switching cost low. Assume any given model could be frozen and plan an exit before you need one.
Third, notice what Amazon kept. It didn't keep the flashy generative models. Reporting says the survivors are the practical, enterprise-facing pieces — a model aimed at building AI agents, and a tool for helping customers customize models to their own data. That tells you where Amazon thinks the money is: not in owning the smartest general model, but in the plumbing that lets businesses actually deploy AI. And that fits Amazon's whole history. AWS didn't win the cloud by having the best individual servers. It won by being the rails everyone else builds on. Amazon has already committed enormous sums to being the compute layer under the leading labs — reporting describes AWS compute agreements worth well over a hundred billion dollars each with the biggest AI companies, and Amazon has put up to eight billion dollars into Anthropic. If those figures wash over you, that's fine. Hold onto one idea instead: Amazon has consistently bet more on being the infrastructure under AI than on winning the model race itself. The frontier model under Abbeel is a hedge against needing anyone else — not necessarily a bet to beat everyone else.
Fourth, the talent story is its own lesson. Reporting noted that Amazon's AGI group ran on a separate pay and leveling system from the rest of the company, specifically to compete for AI researchers. And it still lost its lab founder, its division head, and a dozen key people, and then shut a marquee office. The lesson is not that money doesn't matter — it obviously does. The lesson is that elite AI researchers follow missions and each other more than they follow org charts. When the founder leaves, the people he brought tend to leave too. If you're trying to build or retain an AI team, the org design and the sense of a real mission matter as much as the comp band. A world-class hire embedded in a confused structure is a flight risk, no matter what you're paying.
And then the last thing, the one that's easy to miss under all the corporate noise. Watch what Abbeel actually builds, because it's a tell about where the whole field is going. If the person now running a major frontier model comes from robotics and reinforcement learning, that's a signal that the center of gravity is shifting — away from just scaling up text prediction, toward models that act, that reason through steps, that learn from their own experience rather than from an ever-larger pile of scraped text. The agent era, in other words, and the beginnings of AI that touches the physical world. The reason a robotics professor is a sensible choice to lead a language-model effort is that the line between those two things is starting to disappear.
So when Amazon's frontier model shows up, whatever it ends up being called, don't just read the benchmark scores. Look at how it was trained. Look at whether it learns from doing rather than only from reading. Because the person building it spent twenty years teaching a machine to fly a helicopter upside down by watching a human do it first — and that, more than any single benchmark, is the direction the ground is moving.
sources (69)
- Amazon confirms it's closing key AI site in San Francisco but says work on its top models continues – GeekWire
- Amazon Winds Down Nova Premier, Omni, Reel and Canvas AI Models in Major Strategy Overhaul | MLQ News
- Amazon's AGI Site Closure Signals Shift to Practical AI Solutions | Welcome.AI
- Amazon AI Strategy Explained: Why Amazon Is Dropping Most of Its AI Models - Memeburn
- Amazon Nova AI Models Shift Focus to Frontier Innovation
- Amazon winds down Nova AI models, shifts focus to frontier research
- Amazon Shuts Its AGI Lab and Cuts Jobs to Chase Enterprise AI Instead - Startup Fortune
- Amazon Reshapes AI Strategy After AGI Layoffs, Lab Shutdown
- Amazon Is Gutting Its AI Division After Sustained Failure
- Pieter Abbeel
- Who is Pieter Abbeel? - FourWeekMBA
- Pieter Abbeel - Wikidata
- Pieter Abbeel | The House Fund
- Pieter Abbeel | OpenReview
- Pieter Abbeel | Research UC Berkeley
- Pieter Abbeel – Biography, Net Worth, Career, Companies & Lifestyle [2026]
- Amazon Restructures Its AI Strategy Ahead of Q2 Earnings, Focusing on Frontier Models | KuCoin
- Amazon AI Strategy Reportedly Shifts Before Earnings, Nova AI Models Put In ‘Keep The Lights On’ Mode | Asianet Newsable
- Amazon AI Strategy Reportedly Shifts Before Earnings, Nova AI Models Put In ‘Keep The Lights On’ Mode
- Amazon Kills Most Nova AI Models, Bets Big on a Single Flagship Before Earnings — BigGo Finance
- Amazon is winding down most of its Nova AI models to bet on one frontier model
- Amazon reportedly winds down most Nova models in major AGI strategy overhaul - Neowin
- Amazon just made a major move to reshape its entire AI strategy - TheStreet
- Amazon just made a major move to reshape its entire AI strategy | Louis Velazquez - Official Website, Entrepreneur, Finance, Technology
- Professor Pieter Abbeel - 2021 Summit
- Pieter Abbeel: Pioneering Robot Learning and Artificial Intelligence
- Pieter Abbeel - Amazon Science
- Andrew Ng
- Pieter Abbeel — Grokipedia
- Pieter Abbeel
- Covariant - 2026 Company Profile, Team, Funding & Competitors - Tracxn
- Covariant: Funding, Team & Investors | Startup Intros
- Covariant launches from stealth to bring universal AI to robots
- Pieter Abbeel - Founder, President, Chief Scientist @ Covariant - Crunchbase Person Profile
- Covariant (company)
- Covariant | About
- What is Covariant? Profile, leadership & funding — Komo
- Industrial AI startup Covariant raises a $40M Series B
- Industrial AI startup Covariant raises a $40M Series B
- Pieter Abbeel
- Distinguished Lectures | Department of Computer Science, Columbia University
- Pieter Abbeel - HKU Musketeers Foundation Institute of Data Science
- Pieter Abbeel
- MODEX 2024: Covariant introduces RFM-1 to give robots human-like ability to reason - Robotics 24/7
- MODEX 2024: Covariant introduces RFM-1 to give robots human-like ability to reason - Supply Chain 24/7
- https://covariant.ai/covariant-introduces-rfm-1-to-give-robots-the-human-like-ability-to-reason/
- Covariant Introduces RFM-1 to Give Robots the Human-like Ability to Reason
- Covariant Introduces RFM-1 to Give Robots the Human-like Ability to Reason | Wire | chronicle-tribune.com
- Robots Can Think Like Humans with Covariant's New AI Model
- VAT: Vision Action Transformer by Unlocking Full Representation of ViT
- “Covariant’s RFM-1” – Covariant Build a Language Model that Can Make Robots Act Exactly & Learn like Humans! - DigiAlps LTD
- Covariant RFM-1
- Covariant is building ChatGPT for robots | TechCrunch
- Pieter Abbeel Speaker Bureau & Booking Fees | Aurum Speakers
- Ph.D. Dissertations - Pieter Abbeel
- Pieter Abbeel: Shaping the Future of AI Through His Students - Oreate AI Blog
- Representation Learning for Perception and Control by ...
- Chelsea Finn
- Diversity is All You Need: Learning Skills without a Reward Function
- Sim-to-real transfer of robotic control with dynamics randomization | OpenAI
- Robot Learning in Homes: Improving Generalization and Reducing Dataset Bias
- Towards Space-Based Environmentally-Adaptive Grasping
- Learning dexterity | OpenAI
- DeXtreme: Transfer of Agile In-hand Manipulation from Simulation to Reality
- Time Reversal as Self-Supervision
- ClutterDexGrasp: A Sim-to-Real System for General Dexterous Grasping in Cluttered Scenes
- On the Verge of Solving Rocket League using Deep Reinforcement Learning and Sim-to-sim Transfer
- MDP Playground: An Analysis and Debug Testbed for Reinforcement Learning
- Timeline & Review of OpenAI's Robotic Hand Project | Exxact Blog