IN THE NEWS

How to Deploy Physical AI in Production

Scaling Robotics Podcast host: Vedant Nair

Watch on YouTubeSpotifyApple Podcasts, or wherever else you get your podcasts.

This Episode

In Episode 14, we’re joined by Jeff Mahler, Co-Founder and Chief Technology Officer at Ambi Robotics, for logistics.

https://www.youtube.com/watch?v=CnHI6k1kjQU

Ambi builds AI-powered robots that sort and stack parcels as they move across logistics facilities. Jeff’s team has deployed over 100 robots that have handled more than 150 million packages and 250,000+ operational hours in production, which is over 30 robot years of runtime. And they’ve done it with a team of just 40 people.

Jeff is a skeptic of generalization in physical AI. He thinks the real hard problem is mastery; can we get a robot to do one thing fast, accurately, and reliably enough that a customer like Amazon or FedEx will actually pay for it?

In this episode, we discuss what that takes to pull off. We cover the harnesses that wrap learned models into production systems, build data flywheels, and drive reliability in our robots.

I loved this conversation because it’s clear that Jeff has put in his 10,000 hours with production robots. I’m confident you’ll learn something from this one.

Timestamps

Transcript

Transcript

00:00:00 – Introduction

Vedant Nair 00:00:04

We are back with another episode of Scaling Robotics. This is a series of conversations with leaders at robotics companies who have actually scaled their fleets, and we hope to democratize their hard-earned lessons to benefit the next generation of robotics companies. Today, we have an awesome guest, Jeff from Ambi. Jeff, thanks so much for being here. Do you mind introducing yourself and Ambi Robotics?

Jeff Mahler 00:00:25

Excited to be here, and thanks for having me. I’m Jeff. I’m a co-founder and chief technology officer at Ambi. We’re building AI-powered robots that help people handle more in industrial operations, starting with supply chain. We’re really focused on machines that multiply productivity in tasks like sorting and stacking items as they go from a virtual shopping cart to your front door.

We’ve deployed over 100 robots in our lifetime, over 150 million production packages, over 250,000 operational hours in production, which is over 30 robot years if one robot was running continuously. So we have a lot of experience in the field.

Vedant Nair 00:01:04

And just from the stats alone, it’s clear you guys are one of the few physical AI companies that have really done this at scale. So there’s going to be a lot of interesting things that we can touch on today. You talked about more than 100 robots that you guys have deployed. What’s the scale of the company, scale of the team?

Jeff Mahler 00:01:25

We are about 40 people today. So we’re a very lean team. And we like to focus on multiplying productivity, not just for our customers, but for ourselves. That way we can sustain that many robots out there in production.

00:01:40 – Mastery Over Generalization

Vedant Nair 00:01:40

Definitely want to hear more about how you guys have managed to run so lean at such a scale. Can take us in a different direction, though, to start. One thing that I loved about your content and the thought leadership that you guys have put out at Ambi is your take on building physical AI systems for production use cases, right? Not just for demos, but for real customers, real uptime, real robots. There’s a ton of hype right now about end-to-end robotics models. There’s a lot of really cool demos that people are posting. Maybe first, what’s the most contrarian thing that you believe about physical AI today?

Jeff Mahler 00:02:18

That’s a great question. And I don’t know if this is contrarian or not, but I’ll talk a bit about my perspective on some of the new findings in robots. First of all, I’m super excited about physical AI and everything that’s happening. I think right now where you’re seeing a lot of the popular media coverage is around generalization. There’s this belief that if robots could only generalize to more tasks, then all of a sudden there will be this moment where we unlock commercial production. And in my experience, that’s not true.

The real hard problem in robotics and physical AI, at least on the AI side, is mastery. How do you get a robot to be really, really good at doing one thing very quickly? That means you need a robot to be able to go into a new environment and get to very high levels of throughput, accuracy, uptime, other reliability metrics as fast as possible. And that’s really where a lot of the challenges we’ve seen deploying robots in production are.

Vedant Nair 00:03:22

It sounds like rather than general intelligence, it’s more about specific intelligence. Maybe if you were to steelman why people believe generalization is so important, what do you think that they’re getting wrong? In your experience, specialization is really important. What is the allure of generalization? Why are people so enamored with it, do you think?

Jeff Mahler 00:03:47

A lot of us have gotten into robots through science fiction and our imaginations about what robots can do. And of course, a lot of us imagine things like robot butlers or maybe things like the robots you see in Star Wars, like C-3PO, that can do general tasks. And of course, that’s very exciting. That is something a lot of us are very much working towards in a future state.

But when you look at the reality of the situation and what gets favored in an economic sense, it’s specialization. A customer like a top supply chain company — Amazon, FedEx, UPS — they want to pay for something that is going to optimize their process. They can save a lot of money by saving a fraction of a penny on something. So if they put a robot exactly where a person was that’s kind of good at that task and they can move them around, it’s just not going to be as good as something that is designed for that specific problem.

I’ll give you another example. The Swiss Army knife was invented, at least in the Swiss Army, in the 1800s. But if you were going to build something, you’d probably grab a normal screwdriver, right? Why is that? It’s because you want the right tool for the job. And when you’re doing some kind of work that’s very important to you, you’re always going to prioritize getting the thing that feels right to get that job done.

We see it all around us. If you start looking at our physical world, you’ll notice — why don’t we have one type of car? Why don’t we have one type of keyboard, something that could even be maybe more obviously generalized? People love their specialized keyboards. And then there’s this general economic concept of division of labor we can go back to, which is really just a big capitalist idea, even predating capitalism, about specialization leading to overall growth and productivity. We feel that and see that when we go into our customer environments for sure. Even a machine that is very specialized, if it goes to a new customer and they have slightly different packages or they want to use a slightly different way of sorting, they want to make sure they’re getting the best results with that. So that’s just coming from my experience, but also how I see physical products in the world around us.

00:06:15 – Reliable Intelligence and the AI Harness

Vedant Nair 00:06:15

I love those analogies from the Swiss Army knife and cars. I think that’s super compelling. Part of the specialization that you talked about was not just making the robot intelligent, but making it reliably intelligent, right? Reliable, accurate, have high uptime in production. What is the tangible difference between creating systems that are just intelligent versus reliably intelligent? And what do you see other teams in the industry fall short of when it comes to building reliable intelligence?

Jeff Mahler 00:06:53

It requires putting the focus in different places to get to reliable intelligence. Once you have an AI model that can do a particular task, you’ve got another starting line. That’s great. So now how do you actually get it to providing commercial value to somebody? You need to have things like a harness around that AI. And that’s one of the key things we talk about — things that can plan at a task level, an application level. Fallbacks. If the AI fails to do something accurately, is there another autonomous action we can take? Or do we have an intervention? And intervention systems are, of course, a big part of the industry.

Then there’s safety systems, which are extremely important. How do we make sure people won’t get hurt with these systems? Monitoring systems. Checking for things that could go wrong before they become bigger problems. These predictive analytics are really key to making robots work well in the real world as well.

So getting into that space, it becomes a problem of optimizing the machine for a particular use case with all these variables that are about how to leverage the next best option you can to keep that machine running and at the level of reliability and throughput that the customer expects. That is really what we call Ambi OS — our platform to take physical AI models, post-train them for particular tasks, and then run them inside of this harness where we can get to that customer-expected level of reliability.

I don’t think this is something necessarily that people are getting wrong per se, but I think it’s something that’s undervalued. I’ve come from the academic world. I did my PhD at UC Berkeley. And I know that that world very much prioritizes novelty. And there’s a place for that. It’s very exciting. We never saw a robot fold clothes this well before. We never saw a robot be able to do things like tie knots or peel things off a table before. It’s new. It is super exciting.

But novelty is not what the commercial world is optimizing for. It is optimizing for economic value. At the end of the day, for the most part, customers want to increase their volume somehow, grow their business, maybe save money. But they’re not doing something just purely because it’s new. And that’s one of the shifts that I’ve had to make and my other co-founders who came from the university. I think a lot of folks who today are getting a lot of investment for more of the academic side of things will realize these things as they get into that place of having to do real commercial deployments. But I think the world will be focusing much more on these mastery problems in the next few years because there is that pressure to deploy and create value from all these investments.

Vedant Nair 00:09:56

There’s a few things to pick apart there. But I love what you said about once you build the model, that’s actually the starting point. And a lot of the value to create reliable intelligence is in everything around the model. You mentioned the harness. You mentioned monitoring, safety, fallbacks. Can you talk me through what a harness actually is when you’re deploying a physical AI system? I’m super familiar with the harness in cloud code or these other coding harnesses. But I think most people don’t have a great sense for what a harness is in robotics specifically. I think it’s a very abstract term right now and would love to hear concretely how you think about it and maybe some of the components that actually make up a great harness when deploying physical AI.

Jeff Mahler 00:10:36

Totally. My answer before is largely talking about the software system around the AI. Specifically these models are taking as input usually some observation of the world, mostly in the form of images or videos, and maybe some instructions in the form of language. And they’re outputting something else. A lot of the cutting-edge results are directly outputting actions of the robots. They might also be predicting something like predicting video. There’s been a lot of exciting things there. It could also be a perception system predicting state of the world. But either way, you’ve got a black box that’s taking those inputs and outputs. And the harness is about, first of all, conditioning the inputs that go into the model and taking the outputs and verifying them or maybe even modifying them, so that way they can be safely, reliably, quickly executed on a robot.

I’ll give you a specific example with Prime 1. That is our foundation model we announced about a year ago for the task of picking. What Prime 1 is doing is taking as input a set of images and outputting a pick action, which is an end effector pose, more or less, with maybe a few other variables attached to it. So we need to take that end effector pose, and then we need to actually produce a reliable action from the robot. We’re going to construct the rest of a motion plan, optimize it so it executes very quickly, do a bunch of safety checks, make sure it’s not going to go outside certain bounds, it’s not going to collide with anything. And we may have to throw away some of the actions that have happened along the way to do that.

In addition, if something happens when we execute that plan, we may need to not go back to the AI, but go to a person. Maybe the robot wasn’t able to pick up a package, but maybe it was something completely different, like the robot couldn’t scan barcodes many times in a row. We see all kinds of stuff out there, like damaged barcodes, bad printers that we can’t read barcodes on. We need to detect all these things and be able to have application logic that’s going to snap out of the AI loop and say, hey, person, can you come take a look at this? Kind of the way you would see it happen with cloud code, but I do think there are more layers in the robot harness today to just make all these pieces work.

Vedant Nair 00:13:08

You just said it, right? These models are a black box. My understanding from hearing our customers and who they’ve deployed with is that sometimes that can be problematic. If you’re not exactly sure, if it’s not deterministic but probabilistic what the robot’s going to do, that can be an issue when it comes to deploying something into my production system. It sounds like you’re saying that part of the harness’s role is to make some of these black box decisions more deterministic. And then, like you said, if things are outside the bounds of what is acceptable, let’s remove those from the slate of possibilities of the actions that our robots can take.

Jeff Mahler 00:13:50

Totally. And something I failed to mention on the last answer was that in some cases we might actually completely throw away the AI-generated action and take some other more classical approach. That’s not the best way to handle every single case, but sometimes it is. Sometimes we can look for a flat surface on an item and go suction that or something analogous to it. Because we want to, again, maintain uptime and anything we can do to keep that robot running is going to benefit the customer.

00:14:23 – Benefits of Physical AI in Production

Vedant Nair 00:14:23

You’re talking so much about these layers, this harness that you’ve had to build around physical AI systems. I think we talked a little bit about how some people are a bit too excited about general purpose physical AI. But on the other side of the spectrum, there are folks who are still kind of stuck in the mode of classical controls and classical autonomy. What are the tangible benefits that you’re actually receiving from having physical AI systems? Are you being able to go to new environments faster, handle more edge cases with the types of boxes that you might see? What are those tangible benefits that you’re actually seeing now by deploying these physical AI systems and having to go through building all these harnesses and layers?

Jeff Mahler 00:15:06

Really, it’s all the things you just listed. A big thing for us, really the first unlock kind of getting us to market, was being able to generalize across a variety of different items. That alone was a big unlock at the beginning. We handle parcels in production mostly. So anything that could come to your front door, and that includes things that don’t have a box or a bag around it. We get things like shirts in a bag with a shipping label on them, and we have to be able to handle that.

But going to different environments, things like being able to operate in different lighting conditions, but also handle different types of input and output containers more seamlessly — those are things that would have tripped up a lot of traditional approaches. And in addition to that, we can start to, over time, unlock new kinds of generalization abilities as well. Things like being able to recognize different types of items that are coming through the system. We’ve gotten much better over time at being able to generalize what happens after picking up the item — placing the item, stacking the items. That’s still very much a cutting-edge problem for production physical AI as we see it.

I think as physical AI continues to improve, that horizon is going to keep expanding. There’s going to be a bigger and bigger set of problems that we can solve that just weren’t possible to automate.

00:16:35 – Achieving Mastery and the Data Flywheel

Vedant Nair 00:16:35

The horizon is practically all the work that is here on Earth and beyond. So we all got a lot of work cut out for us after this. But for the tasks that we are focused on now, another thing that you brought up earlier is mastery. I think that makes a lot of sense. You want these really high nines of reliability. You want customers to know that this thing is going to do the job at not just a human skill level, but a superhuman skill level. What are the things that you’re focused on to achieve mastery with any given workflow, any given task? Curious what the strategies are there, the modes of thinking.

Jeff Mahler 00:17:16

First of all, we have to be very thoughtful about what specific applications we even go after in the first place. We get requests all the time from customers to do many different things in logistics. Most of them we say no to. Sometimes it’s related to something about the economics of the problem. But there can sometimes also be an element of, is this problem close enough to what we know our physical AI can do well? And or is it moving in the right direction that we think we need to diversify the data that we’re handling?

Because we are building this data flywheel and improving our models over time. We want them to generalize to more tasks. Stacking is one problem that we decided to go after because we really felt that a big, undervalued piece of the robotics puzzle was generalized placement of items. And I still believe that. Being able to unlock that in production applications can unlock many more things than just building pallets, which is what we do with AmbiStack today.

Coming back to how we actually get to mastery — we have a lot of infrastructure that allows us to do this well. Some of the things that become really important are basically all the data pipelines. We need to know as much as we can about whatever is happening on the robots without maxing out our customers’ bandwidth. Because sometimes that is more limited than you would expect in 2026. But we want to collect all this telemetry data and then make it very accessible so that way we can go back and make very specific changes that are going to improve the robot and have clear rewards for how the robot is improving over time.

So a person could make changes to improve things. The data could be folded back into AI models. Or we can even do things like allow AI to propose some of the changes itself with all the amazing abilities of coding agents today. To give you a more specific example, having infrastructure for doing things like A/B testing — extremely important in robotics. And it is not necessarily the same A/B testing you see across the board on the web. Because you have different robots in different environments handling different packages. It may take different amounts of time for you to get statistically interesting amounts of data on them. And we can’t be doing a lot of exploration in production. We can’t go out there and say, hey, robot, take some random actions and see what happens. My customer will be calling me in the middle of the night and probably yelling at me. I need to make sure that if we are A/B testing something, it is qualified and going to perform at a certain level. So that is a really important system to getting to mastery that we’re very excited about.

Vedant Nair 00:20:24

I keep hearing data pipelines, data flywheel. I think the whole industry is talking about data flywheels. Everyone wants that. And we can see the benefits of what a really large spinning data flywheel gives us. You mentioned some of the things on infrastructure, but zooming out as a whole, what do you think Ambi does really well with building and spinning this data flywheel? One of the things I imagine is that there’s some alpha in actually having these robots deployed in production rather than teleops and test rigs and data collection setups. But I would love to hear about how you think about data flywheels.

Jeff Mahler 00:21:03

Data flywheels — first of all, there’s a lot of debate about them recently. I think part of it is that folks are looking at data flywheels purely from the lens of generalization sometimes. And I don’t think that’s necessarily the right way to think about them. The argument being, production data flywheels might collect a lot of task-specific data, but how are they going to get the diversity of data that takes you to all tasks? And I think that is a legitimate argument. I talked a little bit about how we’re thoughtful about which tasks we go after.

But a really big part of the data flywheel is the mastery piece. As we get to a new task, how fast can we spin up that flywheel and get the robot to where it needs to go? When we deploy a new type of robot, we expect to get there much faster. This is something that we track. When we went AmbiSort — first proof of concept of the AI skill to actually first getting a robot proven in production — it was about four to five years. And then the second time we did that with AmbiStack, getting the base AI capability to actually passing a customer pilot going to production was about a year. So we were able to really close that gap. And we want to do better than that over time too.

Collecting all those edge cases, all that data with hard negative mining about what are all the things that can go wrong when handling these types of items — that’s where we get a reusable benefit with data flywheels. I would say that we at Ambi are very efficient about how we use this data, very targeted and focused on where we want to make improvements for customers. This is partially why we’ve been able to do what we’ve done with such a small team. We can automate a lot of the process of getting that data, looking for what things we should be looking at to get to a better point of mastery. And even now starting to automate some of the next things we should try and A/B testing and so on to get those improvements. So that’s where we’re focused. We’re not the only ones, but I think that’s something we do really well.

00:23:16 – Skill Transfer and the Ambi Skill Suite

Vedant Nair 00:23:16

Can you talk more about what that skill transfer actually looks like? It seems like you’re saying from sort to stack, we were measuring how fast we could go to market. And that time significantly decreased and hopefully each additional skill gets easier and easier. What have you noticed as you have kept building this flywheel? What parts of transferring between skills have gotten easier? What parts remain things that you guys are focused on? Any learnings from this kind of skill transfer are very interesting to me.

Jeff Mahler 00:23:51

A lot of our skills today are around picking and placing items. There’s a lot broader set of robot skills that we could talk about. But what transfers really well? Things that involve item-level knowledge, we’ve been able to get really good at. For example, with picking, we’ve now collected — most of our data is related to some picking tasks, that 150 million sorts that I mentioned earlier. So when we get a new item, we can handle that very quickly. And if we need to pick an item in a different scenario, maybe the item’s not in a bin or a conveyor, it’s in a tote or it’s somewhere else, we can pick that item pretty quickly.

Now, if we move to stacking, we’re dealing with an item that maybe we’ve already grasped and we need to figure out where to put it. Or maybe we have a set of options and we need to think about where to put those items. We have had to do R&D and new development to get ourselves to this stacking-level capability. A lot of it is leveling up to this task-level planning around the sequence of actions that the robot needs to take because our picking systems were not planning for that sequence of actions so much. More or less, if they could pick one or two actions out, they could do really well. So that’s an area where we’ve had to do more development.

I’ll also say we do a lot of state-based AI as well as the harnesses based on state. I think it’s something that a lot of the newer research is moving away from — having a notion of where is this item. But I think it’s really important, and it’s important for a few different reasons. One is safety. If you don’t know where an item is, how do you know that it’s not going to hit somebody? You have to keep eyes on it at all times. Another is just reliability. How do you make sure that if something’s right on the boundary of maybe it can fit in something or not, if you have a way where you can actually measure that, you can get some benefits. A third is integration logic. We do a lot of applications with stack, for instance, where a bunch of items will come down the conveyor. We might have to sort this item to this bin, this item to this bin, this item to this bin. How do you actually do that with an end-to-end model? You can have some sequence of tokens that are the item and where its destination is and so on. But that is, in my mind, a representation of state.

My point is we have a lot of systems that will get some item-level understanding to produce some state representation for the item that can then be reused by other components. And that is one that generalizes quite well between the types of tasks we’ve been working on.

Vedant Nair 00:26:52

What I’m hearing is that this ability to get good at different skills quickly seems to have translated into a product offering. I saw the Ambi Skill Suite. Would love to hear more about that. It sounds like now you guys are even licensing these skills to run on different folks’ hardware. Curious about the decisions behind that and what it’s taken to run these skills on different embodiments or different platforms.

Jeff Mahler 00:27:20

We announced the Skill Suite earlier this year. It was really driven by folks coming to us and looking for our software capability in other problems. For a long time, we’ve been doing these vertically integrated robots as the only business model that we run. But we’ve realized in doing that that there’s a long tail of automation use cases out there, different robots that need to be specialized for different problems. We don’t believe it makes sense for us to go after every single one of them. But if we can provide our skills and add value in that way, we have a way of having our foot in the door on maybe a much broader set of robot applications, which can feed data back to the flywheel, increase revenue, and so on.

We have a number of skills like inspecting items, picking items, placing items, stacking items that we’re now offering to work with partners to take to market. These skills are in some ways abstracted away from the robot embodiment — not in all ways, I’ll come back to that — but they can be used with a supported set of cameras, end effectors, robot arm types, and so on. We’re working now with some initial partners to take the first skills to market in applications.

We got many different inquiries about many different use cases. And again, we’re 40 people. So we said, all right, let’s look at this and figure out what’s the right starting point. We decided to start with a vision problem, which is taking our item analysis and inspection systems and providing that to partners or customers through a partnership with Cognex in order to analyze items before they go into other automation systems. We had developed some unique capabilities around getting out material properties, locating barcodes, being able to read text on shipping labels. Because that’s something that I thought was solved, but in a semi-structured logistics environment, turns out it’s not totally solved. So we’ve got the first deployments of those systems going on right now.

We’re excited about expanding to some other use cases. Particularly, we’re interested in these other types of stacking problems where AmbiStack is our vertically integrated system — it’s a big gantry or Cartesian robot. There’s a lot of applications that are well-serviced by a six-axis arm. So that is what we are looking at as the next phase of AmbiStack’s rollout.

00:30:02 – The Future of Skills and Industry Partnerships

Vedant Nair 00:30:02

Very cool to hear about that partnership with Cognex. Looking forward into the future of where skills takes you or where the industry will be, do you see a day where the big supply chain companies of the world, the FedExes, the Walmarts, are owning the hardware and running the skills themselves? Do you always think there will be automation partners in the loop? I’m just trying to think about analogies from the LLM world where now it feels like enterprises are fine-tuning their own models and doing different post-training for different internal applications. Curious what the parallels will look like for physical AI as the industry matures, as different capabilities come off the shelf like you guys are providing.

Jeff Mahler 00:30:52

There are other trends, one of them is faster and faster shipping. Will we ever get fast enough shipping? It’s an interesting problem. If we invent teleportation, maybe it will become fast enough. But until then, I think there’s an opportunity to keep innovating and improving that part of the experience as well.

My belief is that the big players will take as much in-house as they can. But because things change quickly and they can’t move that quickly, they will continue to work with automation partners who can move faster and help them get a competitive advantage. Maybe at some point there’s some steady state to all that. But that’s hard to picture because things are always changing in unpredictable ways. Supply chain isn’t just about what we buy as consumers either. It’s also about getting stuff to businesses. There’s so much supply chain around semiconductors and batteries and things like that. Who knows what it’ll look like in 10 to 20 years. So I think there will continue to be competition and these big companies needing to work with smaller partners to automate and stay on top.

Vedant Nair 00:32:38

That makes sense. I think there’s probably no such thing as fast enough delivery. I’ve heard the Jeff Bezos quotes — customers are always going to want faster delivery, larger selection, and cheaper prices. We have an insatiable demand for all three of those things. And I would imagine if you guys invented teleportation tomorrow, maybe someone would ask you if your teleportation could become even faster or cheaper. So that makes a ton of sense to me.

00:33:07 – Deploying Robots in Production

Vedant Nair 00:33:07

Moving on to talking about how you actually deploy real robots. We talked a lot about the technology that you guys are building. What are the general principles that you have from the engineering side of what it takes to deploy robots onto the warehouse floor? What are the things that you are prioritizing when you’re thinking about getting a robot to a customer, making it valuable?

Jeff Mahler 00:33:41

There’s a lot to that. A big one is expectation management, which is not entirely an engineering problem. But I want to put it out there because I think not everyone thinks about this — make sure the customer has reasonable expectations for what they’re actually going to get. And the people that are going to run it in the actual building and be responsible for it are embracing it. And their concerns are heard because that’s something that is not always done. It can be easy to overlook if you’re focused on just the engineering problem. But it’s a scoping and communication problem that’s very adjacent to engineering.

When we talk about actually getting into production, a lot of it comes down to testing, trying to make sure that we have done the rigorous testing that is needed to get into production. That includes things like testing for specific edge cases we care about. We have big regression suites of things that we can run offline and make sure we don’t mess up. But the physical testing is something that I don’t see going away. There’s a lot of interest in making this testing faster through things like simulation. I’m all for it. But at the end of the day, if simulation isn’t perfect and I need to be responsible for a robot that’s going to go into a customer environment, then I need to make sure I’ve tested on the physical system.

We will have long test periods where people are just loading items into a robot, watching the outcomes for hours, and then having really good monitoring systems where we can go back. If something weird happened, you can pinpoint where it happened in a video, look at what happened, diagnose those things. For me to feel comfortable with something going onto a customer floor, it’s having done enough cycles of that to hit certain KPIs. Some of them are application specific, like throughput. Some of them are more general, like uptime. We’re going to look at what exceptions came up, how many exceptions there were, and did we keep the robot running as much time as we needed to. Those are the gates we have to pass through.

00:35:53 – Building Testing Environments

Vedant Nair 00:35:53

Testing has always been fascinating to me because some of our customers we see, they have to set up really intricate test environments because you want them to be as close as possible to your customer’s actual environment. So real facilities, like you said, where people are loading boxes or doing different parts of construction or manufacturing. On a very granular level of building testing environments, what are the things that you’re focused on? You want it to be super detailed, but you also can’t literally recreate your customer’s warehouse. What are the things that you’re saying, definitely, yes, we need to cover? And how do you think about those trade-offs when you’re making your testing environments?

Jeff Mahler 00:36:41

One of the things that we prioritize very highly is having very challenging items for the robot to handle. We have historically called these adversarial items, but really it’s just edge case items. Things like a jar of peanut butter in a paper mailer. It behaves very funny — it rolls around and there’s this long strand hanging off of it. So we’ve constructed a lot of these physical items over time that we can go run with. If we’re doing stacking, we might tape a bunch of sheet metal to one side of the box so it’s way off center in terms of where that center of mass is and give that to the robot. That’s one area where we focus a lot of energy because we want to cover the edges of where we think that item distribution could be and what are some of the strange things that could happen as a result of that.

Something that we would not necessarily consider as important in offline testing would be something like we may do a lot of testing without the exact same type or number of containers in AmbiStack. For AmbiStack, for instance, we can put pallets in one area, we can put walled containers in another. And we may not require that we do all of our testing with the entire thing filled with all those containers because we often feel we can learn when we need to.

Something else that’s hard to really get at in the early testing before going to production at all is just the really long tail of things that can happen. It might take millions of cycles before you have a package fly out of the robot’s gripper in a way that damages a camera. We’ll think about that. We’ll theorize about it. We’ll try to guard things. But if that happens randomly once in one of our tests, we may not over-index on it because we know how rare that sort of a thing could be. Of course, we still have to fix the problem. But we have to have some sense of how rare is this problem really. Should this block us moving forward? Or should we maybe have a robust spares program so that way, if this goes down and our support people can get on site in two hours, we can be back up and running really quickly.

Vedant Nair 00:39:05

I really like that in the testing environments, you’re really focused on the edge cases because if you can paint the outer bounds of the possibilities, then you can be a little more sure of what happens inside of it and your robot can handle those things. And also just shout out to all the QA and validation engineers out there. They’re doing God’s work.

00:39:35 – Modular Hardware

Vedant Nair 00:39:35

Another thing I’ve seen Ambi mention when deploying robots is modular hardware. This is really interesting to me. I haven’t seen too many production companies talking about this. Why is modular hardware so important to you guys? Have there been any lessons or anecdotes from your experience that have made you place importance on this kind of hardware?

Jeff Mahler 00:39:53

When we go to a new application, we want to be reaching the levels of mastery as quickly as possible. And one of the best ways to do that is modularity in many things, but hardware is one of them. If we can reuse components that we’ve already really optimized, made robust, and so on, then we’re at a much better starting point when we go into that application.

To give you one specific example — with the AmbiSort robot, it’s taking in a bin of random e-commerce packages. There’s an arm picking them, scanning them, placing them on a platform, and then a cart comes up on a gantry and scoops them and puts them into bags. That gantry piece of the puzzle was something that we initially felt like we just kind of had to do to make the application work and felt kind of bespoke. As we started looking at more customer problems over time, we realized, hey, there’s a big opportunity to reuse this because this is not something you can very easily get off the shelf and get to a high level of reliability.

When we started looking at stacking problems and then we started seeing how many people wanted to sort to many pallets instead of just four-ish around a robot arm, we saw an opportunity to take that same module, extend it to 3D — so it is a new module, but leverage the same software, leverage a lot of the same supply chain, and start at a much higher point of reliability than we did in the past. That’s allowed us to unlock this area where there’s not a lot of other people offering a similar solution, so we have the ability to offer something unique and solve problems that customers couldn’t before. That’s probably the biggest example where we’ve benefited from it, but we also think a lot about this from the perspective of end effectors because those do sometimes need to change between applications, and then some of the other off-the-shelf components with cameras, scanners, and so on.

00:42:04 – Integrating with Other Systems

Vedant Nair 00:42:04

You talked about, for example, in your partnership with Cognex, using vision systems to track objects, maybe classify them before they go into other automation systems that the customer might run. I think this is something that a lot of folks discount — the fact that your system is going to be put in place in production next to not only other humans, but other automated systems as well on the production floor. What have you learned about playing nice with other systems? Is it things you build? Is it how you think about the workflows? What might other folks benefit from focusing on when actually thinking about how their robot is going to work in the broader logistic system that the warehouse they’re deploying to is operating?

Jeff Mahler 00:42:56

Excellent question. One thing is keep the integration as simple as possible. We don’t want our robot to, in order to work, need to get 10 different pieces of information at 10 different times from some upstream system or communicate that much information to a downstream system. Ideally, it’s very small and well-defined touch points between the systems. And that also tends to make it more general as well because more pieces of automation will fit the bill there.

Another big thing we’ve had to think about is rate matching. We go off of sorters a lot of times, and sorters I’m referring to here are conveyors — large conveyor systems where they might be 12 feet high bringing packages across a room and diverting them off down slides into maybe different bins, maybe different conveyors. If we hook up directly to a slide like that and it’s feeding us packages, and we happen to randomly get a burst where there’s many different items that come in and we can’t handle that fast of a burst, then that’s causing a problem for the system as a whole.

We have to think carefully about how we do the right buffering so that way we can rate match our systems. This is why AmbiSort is using this big wheeled blue bin. It seems a little clunky and a lot of times when customers look at it, they’re like, no, I want to connect via conveyor. But when we actually get to deployment, a lot of times we end up with bins and it’s for this reason. You can put the robot in a different part of the building. You can fill up a bunch of bins, queue up the work, run it when you’re ready.

That piece of the puzzle is really key. And that is really fundamental to how we think about robots in industrial operations in general. The process has to be designed in such a way that it can maximally leverage the robot. And that might be different than having a robot being exactly where a person was. You’re flowing a bunch of material or packages or items through a system. You want to process them as efficiently as possible. And that might require a different form factor, a different workflow than what used to be done manually. That is something that should definitely not be discounted when thinking about integrating robots into existing warehouses and automation systems.

00:45:30 – Partnership with Pickle Robot

Vedant Nair 00:45:30

Makes a lot of sense. On that note of integrating with other systems, I saw that you and Pickle Robot had announced a partnership where your two physical AI systems were working in conjunction. Can you tell us a little bit more about that and why that partnership made sense for you guys?

Jeff Mahler 00:45:49

That was spawned from us both being at the same customer and initially being used in two separate areas. Then, of course, there’s the natural, let’s automate this whole thing and put the two systems together. It happened quite fast, actually. The integration, in many ways, was pretty straightforward. They can put an item on a conveyor, it conveys to us, and we pick it up and stack it. We got the whole thing running in maybe just a couple days’ time, which was excellent.

I think it’s partially because we’ve both designed our systems in ways where there’s a standard output flow for them and a standard input flow for us that already kind of matches up. So just running them in their normal state worked really well. I think there’s opportunities to improve some of this, though, particularly around the physical AI portion. Being able to share certain kinds of data, I think, can benefit both systems. That’s something where I think some of the future opportunity is in integrating between these physical systems. That’s something that hasn’t really been done with traditional automation, probably because it’s not really needed and it’s also a lot more of a very closed culture. But that’s something we’re thinking a lot about.

00:47:19 – What’s Next for Ambi

Vedant Nair 00:47:19

That would be really cool and I can see how that would be effective. Are there any other new projects or initiatives at Ambi that you’re excited about that maybe you can talk a little bit more about for the audience?

Jeff Mahler 00:47:34

I’m excited about a lot of different stuff. One of the big things we are really excited to share with the world is that we’ve been able to, since deploying AmbiStack, really push the stacking AI part of our system. That is a new piece of our system which is an AI model that is looking at a number of different action possibilities and the state of the world and looking at a sequence of future actions and deciding what to do because you need to figure out how to play 3D Tetris — put the right item here at this time so that way you can make sure you have a stable stack into the future.

We built a simulation engine where we can simulate this problem. We’ve run reinforcement learning in this simulation to get an initial model that can get into production. And now as we’ve been running with more data from operations, we’ve been starting to get to higher and higher levels of space utilization, getting up to over 70% in actual production environments. There’s still a lot of progress to be made there, but we’re starting to see that performance curve where we know we’ll be able to get to a really good place.

I think it’s a huge productivity opportunity for warehouses where stacking items really densely is kind of hard to do as a person. I’m sure if you have enough time, you can do it well, but if you’re on the clock and you can’t see what items are coming in the future, it’s hard to pay attention to all that. Whereas a robot with AI can pay attention to all the knowledge that it has at one given time in a different sort of way. So I’m really excited about the opportunity for that technology.

Vedant Nair 00:49:25

That will be really cool. And I’m sure you just spoke up a bunch of people who have spent hundreds of hours playing Tetris — like, wow, I need to go work at Ambi.

00:49:40 – Rapid Fire

Vedant Nair 00:49:40

Jeff, this has been an awesome conversation. Thank you so much for your time. To close this out, I would love to do a rapid fire round before we get out of here. How does that sound to you?

Jeff Mahler 00:49:53

Sounds great.

Vedant Nair 00:49:54

What advice would you give to founders or engineers who are about to scale their first real fleet of robots?

Jeff Mahler 00:49:57

My biggest piece of advice is treat operations as a first-class citizen in your organization. I think it’s an afterthought for a lot of robotics companies. That is the face of your company at the customer site. That is how you run the data flywheel. That is how you get to mastery. So put a focus there.

Vedant Nair 00:50:17

Couldn’t agree more. On the flip side, what do you no longer believe about scaling and deploying that you once put a lot of emphasis on?

Jeff Mahler 00:50:25

It’s a lot of what we talked about today. I used to think that capability was really the area to focus and now I think it’s mastery and it’s also integration, which we talked about through the threads of a lot of these questions.

Vedant Nair 00:50:39

You’ve spent basically your academic and professional career, a lot of it, working in robotics. What makes you love robotics?

Jeff Mahler 00:50:44

I love so many elements of robotics. It’s hard to know where to focus there, but of course the sci-fi angle — I was inspired by robots just to get into technology. But I think what’s kept me going in it is the fact that we can create software systems that move the world around us. And what’s really kept me going in the startup world is connecting that technology with people, working with people who don’t care that something is a robot, but you can actually create value for them, remove some of that very injury-prone repetitive motion work that they were doing and level them up into robot operators. I love that element of it.

Vedant Nair 00:51:27

Awesome. And lastly, how can the audience help you? Do you have any shout-outs or call to action? Are you guys hiring? Feel free.

Jeff Mahler 00:51:35

Yes, we are hiring. If what we spoke about today resonates with anyone out there, if you’re in physical AI and you want to focus on connecting that with real commercial problems instead of being back in the lab, then please reach out. We have a number of roles open from software to sales. You can reach me at jeff at ambirobotics.com.

Vedant Nair 00:51:55

Amazing. Jeff, thanks so much for your time. I really enjoyed this conversation. You’re one of the most thoughtful and pragmatic leaders in the space today. And I’m so excited to see what Ambi gets up to in the future. Thanks, everyone, for taking the time.

Jeff Mahler 00:52:10

Thanks, Vedant. Fantastic questions.

Extras

  • Ambi is hiring
  • Brought to you by Miru, RobotOps Infra for Scaling Teams: