A gentle introduction to automated reasoning

Meet Amazon Science’s newest research area.

This week, Amazon Science added automated reasoning to its list of research areas. We made this change because of the impact that automated reasoning is having here at Amazon. For example, Amazon Web Services’ customers now have direct access to automated-reasoning-based features such as IAM Access Analyzer, S3 Block Public Access, or VPC Reachability Analyzer. We also see Amazon development teams integrating automated-reasoning tools into their development processes, raising the bar on the security, durability, availability, and quality of our products.

The goal of this article is to provide a gentle introduction to automated reasoning for the industry professional who knows nothing about the area but is curious to learn more. All you will need to make sense of this article is to be able to read a few small C and Python code fragments. I will refer to a few specialist concepts along the way, but only with the goal of introducing them in an informal manner. I close with links to some of our favorite publicly available tools, videos, books, and articles for those looking to go more in-depth.

Let’s start with a simple example. Consider the following C function:

bool f(unsigned int x, unsigned int y) {
   return (x+y == y+x);
}

Take a few moments to answer the question “Could f ever return false?” This is not a trick question: I’ve purposefully used a simple example to make a point.

To check the answer with exhaustive testing, we could try executing the following doubly nested test loop, which calls f on all possible pairs of values of the type unsigned int:

#include<stdio.h>
#include<stdbool.h>
#include<limits.h>

bool f(unsigned int x, unsigned int y) {
   return (x+y == y+x);
}

void main() {
   for (unsigned int x=0;1;x++) {
      for (unsigned int y=0;1;y++) {
         if (!f(x,y)) printf("Error!\n");
         if (y==UINT_MAX) break;
      }
      if (x==UINT_MAX) break;
   }
}

Unfortunately, even on modern hardware, this doubly nested loop will run for a very long time. I compiled it and ran it on a 2.6 GHz Intel processor for over 48 hours before giving up.

Why does testing take so long? Because UINT_MAX is typically 4,294,967,295, there are 18,446,744,065,119,617,025 separate f calls to consider. On my 2.6 GHz machine, the compiled test loop called f approximately 430 million times a second. But to test all 18 quintillion cases at this performance, we would need over 1,360 years.

When we show the above code to industry professionals, they almost immediately work out that f can't return false as long as the underlying compiler/interpreter and hardware are correct. How do they do that? They reason about it. They remember from their school days that x + y can be rewritten as y + x and conclude that f always returns true.

Re:Invent 2021 keynote address by Peter DeSantis, senior vice president for utility computing at Amazon Web Services
Skip to 15:49 for a discussion of Amazon Web Services' work on automated reasoning.

An automated reasoning tool does this work for us: it attempts to answer questions about a program (or a logic formula) by using known techniques from mathematics. In this case, the tool would use algebra to deduce that x + y == y + x can be replaced with the simple expression true.

Automated-reasoning tools can be incredibly fast, even when the domains are infinite (e.g., unbounded mathematical integers rather than finite C ints). Unfortunately, the tools may answer “Don’t know” in some instances. We'll see a famous example of that below.

The science of automated reasoning is essentially focused on driving the frequency of these “Don’t know” answers down as far as possible: the less often the tools report "Don't know" (or time out while trying), the more useful they are.

Today’s tools are able to give answers for programs and queries where yesterday’s tools could not. Tomorrow’s tools will be even more powerful. We are seeing rapid progress in this field, which is why at Amazon, we are increasingly getting so much value from it. In fact, we see automated reasoning forming its own Amazon-style virtuous cycle, where more input problems to our tools drive improvements to the tools, which encourages more use of the tools.

A slightly more complex example. Now that we know the rough outlines of what automated reasoning is, the next small example gives a slightly more realistic taste of the sort of complexity that the tools are managing for us.

void g(int x, int y) {
   if (y > 0)
      while (x > y)
         x = x - y;
}

Or, alternatively, consider a similar Python program over unbounded integers:

def g(x, y):
   assert isinstance(x, int) and isinstance(y, int)
   if y > 0:
      while x > y:
         x = x - y

Try to answer this question: “Does g always eventually return control back to its caller?”

When we show this program to industry professionals, they usually figure out the right answer quickly. A few, especially those who are aware of results in theoretical computer science, sometimes mistakenly think that we can't answer this question, with the rationale “This is an example of the halting problem, which has been proved insoluble”. In fact, we can reason about the halting behavior for specific programs, including this one. We’ll talk more about that later.

Here’s the reasoning that most industry professionals use when looking at this problem:

  1. In the case where y is not positive, execution jumps to the end of the function g. That’s the easy case.
  2. If, in every iteration of the loop, the value of the variable x decreases, then eventually, the loop condition x > y will fail, and the end of g will be reached.
  3. The value of x always decreases only if y is always positive, because only then does the update to x (i.e., x = x - y) decrease x. But y’s positivity is established by the conditional expression, so x always decreases.

The experienced programmer will usually worry about underflow in the x = x - y command of the C program but will then notice that x > y before the update to x and thus cannot underflow.

If you carried out the three steps above yourself, you now have a very intuitive view of the type of thinking an automated-reasoning tool is performing on our behalf when reasoning about a computer program. There are many nitty-gritty details that the tools have to face (e.g., heaps, stacks, strings, pointer arithmetic, recursion, concurrency, callbacks, etc.), but there’s also decades of research papers on techniques for handling these and other topics, along with various practical tools that put these ideas to work.

Policy-code.gif
Automated reasoning can be applied to both policies (top) and code (bottom). In both cases, an essential step is reasoning about what's always true.

The main takeaway is that automated-reasoning tools are usually working through the three steps above on our behalf: Item 1 is reasoning about the program’s control structure. Item 2 is reasoning about what is eventually true within the program. Item 3 is reasoning about what is always true in the program.

Note that configuration artifacts such as AWS resource policies, VPC network descriptions, or even makefiles can be thought of as code. This viewpoint allows us to use the same techniques we use to reason about C or Python code to answer questions about the interpretation of configurations. It’s this insight that gives us tools like IAM Access Analyzer or VPC Reachability Analyzer.

An end to testing?

As we saw above when looking at f and g, automated reasoning can be dramatically faster than exhaustive testing. With tools available today, we can show properties of f or g in milliseconds, rather than waiting lifetimes with exhaustive testing.

Can we throw away our testing tools now and just move to automated reasoning? Not quite. Yes, we can dramatically reduce our dependency on testing, but we will not be completely eliminating it any time soon, if ever. Consider our first example:

bool f(unsigned int x, unsigned int y) {
   return (x + y == y + x);
}

Recall the worry that a buggy compiler or microprocessor could in fact cause an executable program constructed from this source code to return false. We might also need to worry about the language runtime. For example, the C math library or the Python garbage collector might have bugs that cause a program to misbehave.

What’s interesting about testing, and something we often forget, is that it’s doing much more than just telling us about the C or Python source code. It’s also testing the compiler, the runtime, the interpreter, the microprocessor, etc. A test failure could be rooted in any of those tools in the stack.

Automated reasoning, in contrast, is usually applied to just one layer of that stack — the source code itself, or sometimes the compiler or the microprocessor. What we find so valuable about reasoning is it allows us to clearly define both what we do know and what we do not know about the layer under inspection.

Furthermore, the models of the surrounding environment (e.g., the compiler or the procedure calling our procedure) used by the automated-reasoning tool make our assumptions very precise. Separating the layers of the computational stack helps make better use of our time, energy, and money and the capabilities of the tools today and tomorrow.

Unfortunately, we will almost always need to make assumptions about something when using automated reasoning — for example, the principles of physics that govern our silicon chips. Thus, testing will never be fully replaced. We will want to perform end-to-end testing to try and validate our assumptions as best we can.

An impossible program

I previously mentioned that automated-reasoning tools sometimes return “Don’t know” rather than “yes” or “no”. They also sometimes run forever (or time out), thus never returning an answer. Let’s look at the famous "halting problem" program, in which we know tools cannot return “yes” or “no”.

Imagine that we have an automated-reasoning API, called terminates, that returns “yes” if a C function always terminates or “no” when the function could execute forever. As an example, we could build such an API using the tool described here (shameless self-promotion of author’s previous work). To get the idea of what a termination tool can do for us, consider two basic C functions, g (from above),

void g(int x, int y) {
   if (y > 0)
      while (x > y)
         x = x - y;
}

and g2:

void g2(int x, int y) {
   while (x > y)
      x = x - y;
}

For the reasons we have already discussed, the function g always returns control back to its caller, so terminates(g) should return true. Meanwhile, terminates(g2) should return false because, for example, g2(5, 0) will never terminate.

Now comes the difficult function. Consider h:

void h() {
   if terminates(h) while(1){}
}

Notice that it's recursive. What’s the right answer for terminates(h)? The answer cannot be "yes". It also cannot be "no". Why?

Imagine that terminates(h) were to return "yes". If you read the code of h, you’ll see that in this case, the function does not terminate because of the conditional statement in the code of h that will execute the infinite loop while(1){}. Thus, in this case, the terminates(h) answer would be wrong, because h is defined recursively, calling terminates on itself.

Similarly, if terminates(h) were to return "no", then h would in fact terminate and return control to its caller, because the if case of the conditional statement is not met, and there is no else branch. Again, the answer would be wrong. This is why the “Don’t know” answer is actually unavoidable in this case.

The program h is a variation of examples given in Turing’s famous 1936 paper on decidability and Gödel’s incompleteness theorems from 1931. These papers tell us that problems like the halting problem cannot be “solved”, if bysolved” we mean that the solution procedure itself always terminates and answers either “yes” or “no” but never “Don’t know”. But that is not the definition of “solved” that many of us have in mind. For many of us, a tool that sometimes times out or occasionally returns “Don’t know” but, when it gives an answer, always gives the right answer is good enough.

This problem is analogous to airline travel: we know it’s not 100% safe, because crashes have happened in the past, and we are sure that they will happen in the future. But when you land safely, you know it worked that time. The goal of the airline industry is to reduce failure as much as possible, even though it’s in principle unavoidable.

To put that in the context of automated reasoning: for some programs, like h, we can never improve the tool enough to replace the "Don't know" answer. But there are many other cases where today's tools answer "Don't know", but future tools may be able to answer "yes" or "no". The modern scientific challenge for automated-reasoning subject-matter experts is to get the practical tools to return “yes” or “no” as often as possible. As an example of current work, check out CMU professor and Amazon Scholar Marijn Heule and his quest to solve the Collatz termination problem.

Another thing to keep in mind is that automated-reasoning tools are regularly trying to solve “intractable” problems, e.g., problems in the NP complexity class. Here, the same thinking applies that we saw in the case of the halting problem: automated-reasoning tools have powerful heuristics that often work around the intractability problem for specific cases, but those heuristics can (and sometimes do) fail, resulting in “Don’t know” answers or impractically long execution time. The science is to improve the heuristics to minimize that problem.

Nomenclature

A host of names are used in the scientific literature to describe interrelated topics, of which automated reasoning is just one. Here’s a quick glossary:

  • logic is a formal and mechanical system for defining what is true and untrue. Examples: propositional logic or first-order logic.
  • theorem is a true statement in logic. Example: the four-color theorem.
  • proof is a valid argument in logic of a theorem. Example: Gonthier's proof of the four-color theorem
  • mechanical theorem prover is a semi-automated-reasoning tool that checks a machine-readable expression of a proof often written down by a human. These tools often require human guidance. Example: HOL-light, from Amazon researcher John Harrison
  • Formal verification is the use of theorem proving when applied to models of computer systems to prove desired properties of the systems. Example: the CompCert verified C compiler
  • Formal methods is the broadest term, meaning simply the use of logic to reason formally about models of systems. 
  • Automated reasoning focuses on the automation of formal methods. 
  • semi-automated-reasoning tool is one that requires hints from the user but still finds valid proofs in logic. 

As you can see, we have a choice of monikers when working in this space. At Amazon, we’ve chosen to use automated reasoning, as we think it best captures our ambition for automation and scale. In practice, some of our internal teams use both automated and semi-automated reasoning tools, because the scientists we've hired can often get semi-automated reasoning tools to succeed where the heuristics in fully automated reasoning might fail. For our externally facing customer features, we currently use only fully automated approaches.

Next steps

In this essay, I’ve introduced the idea of automated reasoning, with the smallest of toy programs. I haven’t described how to handle realistic programs, with heap or concurrency. In fact, there are a wide variety of automated-reasoning tools and techniques, solving problems in all kinds of different domains, some of them quite narrow. To describe them all and the many branches and sub-disciplines of the field (e.g. “SMT solving”, “higher-order logic theorem proving”, “separation logic”) would take thousands of blogs posts and books.

Automated reasoning goes back to the early inventors of computers. And logic itself (which automated reasoning attempts to solve) is thousands of years old. In order to keep this post brief, I’ll stop here and suggest further reading. Note that it’s very easy to get lost in the weeds reading depth-first into this area, and you could emerge more confused than when you started. I encourage you to use a bounded depth-first search approach, looking sequentially at a wide variety of tools and techniques in only some detail and then moving on, rather than learning only one aspect deeply.

Suggested books:

International conferences/workshops:

Tool competitions:

Some tools:

Interviews of Amazon staff about their use of automated reasoning:

AWS Lectures aimed at customers and industry:

AWS talks aimed at the automated-reasoning science community:

AWS blog posts and informational videos:

Some course notes by Amazon Scholars who are also university professors:

A fun deep track:

Some algorithms found in the automated theorem provers we use today date as far back as 1959, when Hao Wang used automated reasoning to prove the theorems from Principia Mathematica.

Research areas

Related content

IN, KA, Bengaluru
Every product a customer returns is a moment where Amazon either recovers value or writes it off — and India's ReCommerce business is on a multi-million-dollar mission to recover more of it, more intelligently, at scale. Machine learning is the core lever: predicting whether a returned unit is sellable without a human touching it, detecting damage and fraud inside sealed packaging from images, routing each unit to its highest-value disposition, and pricing recovered inventory dynamically. India's returns network is large, fast-growing, and structurally different from other geographies — a rich, high-impact environment for an Applied Scientist to build models that move real financial and customer-experience metrics. We are hiring an Applied Scientist to build and adapt the ML that powers India ReCommerce. You will work at the intersection of two mandates: building India-first models for problems unique to our market, and adapting proven Worldwide models to India's data, catalog, and operational reality — recalibrating them where distribution, language, and process differ. You will own problems end-to-end, from framing and data through modeling, evaluation, and production deployment, partnering closely with engineering, product, and operations. Key job responsibilities Build ML models for automated returns grading — predicting the salability of returned units from structured and unstructured signals so units can be evaluated with zero or minimal human touch, improving speed, accuracy, and recovery value. Develop computer-vision models for defect detection, condition assessment, and anomaly/fraud identification (including inside sealed packaging), and for establishing chain-of-custody and damage attribution across the returns journey. Build disposition-prediction and routing models that direct each unit to its highest-value recovery path (resale, repair, liquidation, donation, recycle) as early as possible in the network. Develop pricing and recovery-optimization models for liquidation and resale, moving from flat rates toward dynamic, grade- and condition-aware pricing. Adapt Worldwide ML models to India — retraining, recalibrating, and re-evaluating for India's return distribution, catalog, languages, and operational constraints, and closing the gaps that prevent a direct lift-and-shift. Own the full model lifecycle — problem framing, data pipelines, feature engineering, training, offline/online evaluation, monitoring, and retraining — with rigorous attention to calibration, drift, and business-metric impact. Partner cross-functionally with engineering (to productionize), product (to frame problems and measure impact), and operations (to ground models in how the network actually runs), and use modern GenAI/LLM tooling to accelerate research and delivery. A day in the life You start by reviewing the performance of a grading model in production — checking calibration and drift against last week's returns, and confirming the recovery-value lift is holding. Mid-morning, you dig into a computer-vision problem: improving detection of a damage type that's driving write-offs, using images captured across the returns journey. In the afternoon you work with a Worldwide science team to bring one of their models to India — scoping what retraining and recalibration India's data requires — then pair with an engineer to move your latest model toward production behind a clean evaluation gate. You close by framing a new problem with a product partner: quantifying the opportunity, defining the label and success metric, and sketching the modeling approach. About the team India ReCommerce owns the systems and science that turn returned and unsellable inventory into recovered value and a better customer experience. You will join a team building an increasingly automated, ML-driven returns network — leveraging Worldwide platforms where they fit and building India-first capabilities where they don't. It is a high-ownership environment with a direct line from your models to measurable business and customer outcomes.
US, VA, Arlington
How do you measure what makes a great leader? How do you evaluate a development program when outcomes take years to materialize and clean experimental conditions are rarely available? How do you take a scientific methodology that a researcher validated carefully in one context and turn it into a system that any HR team across a company of over a million employees can run on their own? These are the kinds of questions the Senior Talent and Transformation Science team works on inside Amazon's People eXperience and Technology organization, and they are questions that matter: the systems this team builds shape how Amazon identifies, develops, and invests in its most senior leaders. As an Applied Scientist on this team you are the person who closes the gap between a validated scientific methodology and a system that runs in production without a scientist standing next to it. The architectural decisions about how scientific methods get encoded into software, the engineering quality bar for the code that implements them, and the reliability of the pipelines that other teams depend on are yours to own. You will work alongside Senior and Principal Research Scientists, an Amazon Scholar, Product Management, and a Senior Applied Scientist who bring deep expertise in behavioral science, psychometrics, and causal inference, and you will be the driving force behind turning that expertise into working, deployable systems for our Amazon executives. The problems you will be building for are genuinely hard and largely unsolved. Scoring a simulation-based leadership assessment with an LLM requires both measurement rigor and a production system that behaves consistently at scale. Estimating the effect of a talent program on leader outcomes requires both a defensible identification strategy and an analytical pipeline someone else can run and trust. Building a self-serve tool that lets a PXT team evaluate a new feature without calling a scientist requires both sound methodology and software that is robust enough to operate without expert supervision. If you want to do work that is technically demanding, scientifically cutting edge, and consequential for real leaders in a large organization, this is that role. Key job responsibilities • Own the production implementation of the team's scientific systems from end to end. When the team validates a new assessment methodology, evaluation framework, or causal identification strategy, you are the scientist who translates it into code that runs reliably, scales, and does not require a scientist standing next to it to operate. • Make the architectural and tooling decisions that determine how scientific methods get encoded into software on this team, choosing abstractions, data structures, and system designs that make the team's scientific components testable, maintainable, and extensible over time. • Define and hold the engineering quality bar for scientific code across the team, establishing and modeling best practices for testing, documentation, reproducibility, and peer review of code in a research team that does not have dedicated software development engineers. • Build the LLM-powered pipelines that operationalize the team's people science, including prompt orchestration, retrieval grounding, automated scoring, and LLM-as-judge evaluation harnesses, writing the implementation yourself and owning the quality and reliability of those systems once deployed. • Extend and adapt scientific techniques at the product level when established approaches fall short. When scoring a simulation-based assessment, estimating a program effect under unusual identification constraints, or evaluating a novel AI feature requires a methodological contribution that does not yet exist, you devise and implement that solution. • Partner with the Research Scientists during methodology design to surface implementation feasibility and trade-offs early, contributing your own scientific judgment on what can be built rigorously within real production constraints before design decisions become expensive to reverse. • Build reusable scientific components, services, and templates that encode methodology once and allow downstream teams to run it without scientist involvement, making the team's research operational infrastructure rather than a bespoke consulting engagement. • Contribute to the design and execution of quasi-experimental evaluations of people programs, owning the analytical implementation and the code pipelines that produce defensible causal evidence from observational and field data. • Mentor scientists on the team on software engineering practices and applied implementation, and participate actively in peer review of experiment designs, analytical approaches, and scientific code written by others. • Communicate implementation trade-offs and system design decisions clearly to product and HR partners in written documents that connect technical choices to business outcomes. A day in the life Your day is anchored in building and testing. You might spend the morning working through a thorny implementation problem, figuring out how to encode a psychometric scoring model into a pipeline that holds up under the messiness of real production data, debugging an LLM evaluation harness that is behaving inconsistently across assessment scenarios, or refactoring a causal estimation component so that another team can run it without calling you first. In the afternoon a Research Scientist might pull you into a methodology design conversation, and your job in that room is not just to follow along but to push back on approaches that would be difficult or brittle to implement, and to propose alternatives that preserve scientific rigor while actually being buildable. You might then shift to reviewing a colleague's code, writing documentation that makes a deployed pipeline understandable to someone who was not in the room when it was designed, or working through a data pipeline problem that is blocking the team's ability to evaluate a new product feature. At the end of most days something that was not working is now working, and the science the team does is a little more durable and a little more independent of any one person than it was in the morning.
US, CA, Culver City
Prime Video is an industry leading, high-growth business and a critical driver of Amazon Prime subscriptions, which contributes to customer loyalty and lifetime value. Prime Video is a digital video streaming and download service that offers Amazon customers the ability to rent, purchase or subscribe to a huge catalog of videos. In addition, Prime Video offers a variety of live sport streaming services in multiple locales. The Prime Video Economist team is looking for an Economist to support PV content valuation. As an economist focusing on Prime Video, you will be responsible for understanding the value that the business creates for our customers and to develop new, disruptive innovations to grow global Prime Video usage and customer value. This role requires an individual with strong quantitative modeling skills and the ability to apply statistical/machine learning, structural models, and experimental design methods to large amount of individual level data. The candidate should have strong communication skills, be able to work closely with stakeholders and translate data-driven findings into actionable insights. The successful candidate will be a self-starter comfortable with ambiguity, with strong attention to detail and ability to work in a fast-paced and ever-changing environment. Key job responsibilities The candidate's responsibilities will include: - Build scalable analytic solutions using state of the art tools based on large datasets - Build causal inference models, conduct statistical/machine learning analyses, or design experiments to measure the value of the business and its many features - Partner closely with Business, Finance, Science, and Tech partners to build prototypes and implement production solutions - Independently identify new opportunities for leveraging economic insights and models in the Video business - Develop and execute product workplans from concept, prototype to production incorporating feedback from customers, scientists and business leaders - Write both technical white papers and business-facing documents to clearly explain complex technical concepts to audiences with diverse business/scientific backgrounds
US, NY, New York
We are seeking a Human-Robot Interaction (HRI) Applied Scientist to develop cutting-edge interactions that make robots feel alive, personal, and fun. In this role, you will focus on verbal and non-verbal conversational systems, social dynamics, memory, and long-term relationship formation between robots, their environments, and the people they interact with. Your contributions will be essential in advancing robotics by enabling expressive, socially intelligent, and trustworthy interactions between robots and humans. Key job responsibilities - Develop interactive systems that leverage large language models, multimodal inputs and outputs, reinforcement learning from human feedback, or other advanced techniques to achieve fluid, engaging, and socially appropriate robot behavior - Design and implement intelligent conversational systems that handle turn-taking, grounding, interruption, and incorporates context drawn from a robot's physical environment and shared history with a user - Integrate perceptual sensor streams including gaze, facial expression, gesture, posture, and more to understand social context and produce coherent, lifelike interactions. - Develop memory and personalization systems that allow robots to form lasting relationships with individual users, learn their environments, and adapt their behavior over weeks and months - Stay updated on advancements in HRI, NLP, multimodal AI, and cognitive and social science to apply cutting-edge techniques to robot interaction challenges - Lead technical projects from conception through production deployment - Mentor junior scientists and engineers - Bridge research initiatives with practical engineering implementation
US, CA, San Francisco
Amazon is on a mission to redefine the future of automation — and we're looking for exceptional talent to help lead the way. We are building the next generation of advanced robotic systems that seamlessly blend cutting-edge AI, sophisticated control systems, and novel mechanical design to create adaptable, intelligent automation solutions capable of operating safely alongside humans in dynamic, real-world environments. At Amazon, we leverage the power of machine learning, artificial intelligence, and advanced robotics to solve some of the most complex operational challenges at a scale unlike anywhere else in the world. Our fleet of robots spans hundreds of facilities globally, working in sophisticated coordination to deliver on our promise of customer excellence — and we're just getting started. As a Scientist in Robot Navigation, you will be at the forefront of this transformation — architecting and delivering navigation systems that are intelligent, safe, and scalable. You will bring deep expertise in learning-based planning and control, a strong understanding of foundation models and their application to embodied agents, and as well as have in-depth understanding of control-theoretic approaches such as model predictive control (MPC)-based trajectory planning. You will develop navigation solutions that seamlessly blend data-driven intelligence with principled control-theoretic guarantees. Our vision is bold: to build navigation systems that allow robots to move fluidly and safely through dynamic environments — understanding context, anticipating change, and adapting in real time. You will lead research that bridges the gap between cutting-edge academic advances and production grade deployment, collaborating with world-class teams pushing the boundaries of robotic autonomy, manipulation, and human-robot interaction. Join us in building the next generation of intelligent navigation systems that will define the future of autonomous robotics at scale. Key job responsibilities - Design, develop, and deploy perception algorithms for robotics systems, including object detection, segmentation, tracking, depth estimation, and scene understanding - Lead research initiatives in computer vision, sensor fusion and 3D perception - Collaborate with cross-functional teams including robotics engineers, software engineers, and product managers to define and deliver perception capabilities - Drive end-to-end ownership of ML models — from data collection and labeling strategy to training, evaluation, and deployment - Mentor junior scientists and engineers; contribute to a culture of technical excellence - Define and track key metrics to measure perception system performance in real-world environments - Publish research findings in top-tier venues (CVPR, ICCV, ECCV, ICRA, NeurIPS, etc.) and contribute to patents A day in the life - Train ML models for deployment in simulation and real-world robots, identify and document their limitations post-deployment - Drive technical discussions within your team and with key stakeholders to develop innovative solutions to address identified limitations - Actively contribute to brainstorming sessions on adjacent topics, bringing fresh perspectives that help peers grow and succeed — and in doing so, build lasting trust across the team - Mentor team members while maintaining significant hands-on contribution to technical solutions About the team Our team is a group is a diverse group of scientists and engineers passionate about building intelligent machines. We value curiosity, rigor, and a bias for action. We believe in learning from failure and iterating quickly toward solutions that matter.
US, NY, New York
We are seeking a Human-Robot Interaction (HRI) Applied Scientist to develop cutting-edge interactions that make robots feel alive, personal, and fun. In this role, you will focus on verbal and non-verbal conversational systems, social dynamics, memory, and long-term relationship formation between robots, their environments, and the people they interact with. Your contributions will be essential in advancing robotics by enabling expressive, socially intelligent, and trustworthy interactions between robots and humans. Key job responsibilities - Develop interactive systems that leverage large language models, multimodal inputs and outputs, reinforcement learning from human feedback, or other advanced techniques to achieve fluid, engaging, and socially appropriate robot behavior - Design and implement intelligent conversational systems that handle turn-taking, grounding, interruption, and incorporates context drawn from a robot's physical environment and shared history with a user - Integrate perceptual sensor streams including gaze, facial expression, gesture, posture, and more to understand social context and produce coherent, lifelike interactions. - Develop memory and personalization systems that allow robots to form lasting relationships with individual users, learn their environments, and adapt their behavior over weeks and months - Stay updated on advancements in HRI, NLP, multimodal AI, and cognitive and social science to apply cutting-edge techniques to robot interaction challenges - Lead technical projects from conception through production deployment - Mentor junior scientists and engineers - Bridge research initiatives with practical engineering implementation
US, WA, Seattle
Our organization in Amazon Robotics builds robots that perform contact-rich manipulation safely and reliably in complex, unstructured environments, at Amazon scale. Our scientists and engineers push the boundaries of robotic manipulation to handle enormous object diversity, bringing deep expertise across planning, control, perception, and machine learning. We learn from real-world data at a scale that few teams in robotics can access. We are seeking an experienced Senior Applied Scientist to help guide a small team advancing reinforcement learning for manipulation. We are creating robots that learn how to push, flip, rearrange, and dexterously insert items with unparalleled robustness, speed, and reliability. Our goal is to deploy robots that will work across Amazon's global network and can handle the full diversity of items that Amazon sells. You will set the technical direction for how we learn these behaviors, from simulation training through reliable execution on physical robots, and you will demonstrate new manipulation capabilities on real hardware at scale. This team's mission reaches beyond any single product: to invent and apply manipulation capabilities that generalize to many future robotics applications. The robots our organization already deploys at scale give you a rare proving ground to collect data, run experiments, and get new policies onto real hardware faster than almost anywhere in the field. You will raise the bar for scientific rigor and engineering quality, and mentor other scientists as the team grows. Key job responsibilities - Set the technical direction for learning non-prehensile and contact-rich manipulation policies, from testing the latest advances in the field through demonstrated capability on hardware. - Oversee the development of reinforcement learning approaches that address the long tail of diverse, demanding manipulation conditions. - Own the path from simulation training to reliable, real-time execution on physical robots, making evidence based calls on where learned approaches should replace engineered ones. - Demonstrate new manipulation capabilities on real robots at scale, and turn one-off results into repeatable methods. - Establish the standards, evaluation practices, and data-informed improvement loops that the team builds on. - Mentor scientists and engineers, and raise the bar for applied science rigor and engineering quality. - Partner across control, perception, and hardware to integrate learned behaviors into working systems. - Represent Amazon in academia through publications and scientific presentations. A day in the life Amazon offers a full range of benefits that support you and eligible family members, including domestic partners. Benefits can vary by location, the number of regularly scheduled hours you work, length of employment, and job status such as seasonal or temporary employment. The benefits that generally apply to regular, full-time employees include: 1. Medical, Dental, and Vision Coverage 2. Maternity and Parental Leave Options 3. Paid Time Off (PTO) 4. 401(k) Plan If you are not sure that every qualification on the list above describes you exactly, we'd still love to hear from you! At Amazon, we value people with unique backgrounds, experiences, and skillsets. If you’re passionate about this role and want to make an impact on a global scale, please apply
IN, MH, Mumbai
Amazon Science gives you insight into the company’s approach to customer-obsessed scientific innovation. Amazon fundamentally believes that scientific innovation is essential to being the most customer-centric company in the world. It’s the company’s ability to have an impact at scale that allows us to attract some of the brightest minds in artificial intelligence and related fields. Our scientists continue to publish, teach, and engage with the academic community, in addition to utilizing our working backwards method to enrich the way we live and work. Please visit https://www.amazon.science for more information. About Amazon Prime Video “Many of the problems we face have no textbook solution, and so we-happily-invent new ones.” – Jeff Bezos
 The Amazon Prime Video team is shaping the future of digital video entertainment. We are seeking a Data Scientist to uncover key insights on how consumers watch videos on Amazon. The ideal candidate will be an expert in the areas of data science, machine learning and statistics, having hands-on experience with multiple improvement initiatives as well as balancing technical and business judgment to make the right decisions about technology, models and methodologies. As consumers increasingly consume digital video, we need to make agile decisions based on what content appeals to our customers. As a Data Scientist at Amazon Prime Video APAC and ANZ analytics team, you will have the opportunity to work on one of the world's largest consumer data sets, influence the long term evolution of our analytics capability and support the expansion of Amazon's digital video business. The Data Scientist will work closely with other research scientists, machine-learning experts, and economists to design and run experiments, research new algorithms, and find new ways to improve optimization across all our associate facing tools. 
 A successful candidate will be able to understand and manage key operational and technical concepts. They will have excellent project and communication skills, and motivation to achieve results in a fast-paced environment. Candidates should demonstrate a passion for working on behalf of customers, have a record of accomplishment of timely delivery of large-scale projects, and have the ability to influence multiple global teams. Autonomy, judgment, influence, and leadership skills are essential. This person will be responsible for ensuring we meet our key deliverables, on time with high quality, and communicating status to internal and external stakeholders. Key Responsibilities - Support the Content team on business reporting, ad hoc analysis, statistical inference and predictive modelling for all Prime Video APAC and ANZ. - Mine and analyze data pertaining to customers viewing experiences to identify critical business insight and make recommendations to optimize content selection. - Proactively develop new ML models using streaming, video, audio and textual data to understand and predict customer streaming behaviour - Translate analytic insights into concrete, actionable recommendations for business or product improvement. Develop and present these as papers to senior stakeholders. - Liaise with your peers in other prime video territories to develop solutions that greatly benefit our global customers - This role will be based in Mumbai, India
US, WA, Seattle
Join us at the forefront of Amazon's sustainability initiatives to work on environmental and social advancements that support Amazon's long-term worldwide sustainability strategy. At Amazon, we're working to be the most customer-centric company on earth. To get there, we need exceptionally talented, bright, and driven people. We are looking for a Senior Research Scientist to join our growing Sustainability team to drive the science behind value chain decarbonization. This role will establish Amazon's scientific methodologies for sector- and cross-sectoral decarbonization mechanisms and establish benchmarks for automated validation and risk assessment. As a Senior Research Scientist, you will be responsible for independently leading assessments of environmental issues across the full spectrum of Amazon businesses and evaluating sustainability impacts across the value chain. You will independently develop quality frameworks and methodologies that enable Amazon to scale procurement of high-quality environmental interventions while maintaining scientific rigor and environmental integrity. Key job responsibilities - Develop quality assessment frameworks for complex environmental interventions, baseline-setting approaches, and measurement methodologies - Build quantitative benchmark and statistical models that enable scalable evaluation across heterogeneous data sources - Create attribution methodologies for supply chain interventions across Amazon's diverse footprint - Develop social and environmental safeguard criteria that integrate community impact assessments - Collaborate with cross-functional teams including procurement, sustainability operations, and business units to translate scientific methodologies into operational requirements - Work under the direction of senior business leaders while acting as lead Subject Matter Expert for value chain decarbonization science, including designing and leading research, data collection, modeling, documentation, interpretation, and validation About the team Diverse Experiences: Worldwide Sustainability values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying. Inclusive Team Culture: It’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (inclusive diversity) conferences, inspire us to never stop embracing our uniqueness. Mentorship & Career Growth: We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.
CN, 31, Shanghai
Worldwide Global Selling has been helping individuals and businesses increase sales and reach new customers around the globe. Today, more than 50% of Amazon's total unit sales come from third-party selection. The Global Selling team in China is responsible for recruiting local businesses to sell on Amazon's 19+ overseas marketplaces and supporting local Sellers' success and growth on Amazon. Our vision is to be the first choice for all types of Chinese business to go globally. The Worldwide Global Selling Analytics, Intelligence, and Technology (WWGS-AIT) team serves as the research, automation, and insight arm of the International Seller Service data hub, enabling rapid delivery of growth insights through strategic investments in regional data foundations, self-service business intelligence solutions, and artificial intelligence tools. The WWGS-AIT team is positioned to establish AI-ready foundational capabilities across the WWGS organization while maintaining excellence in business insight generation, and self-service BI/AI application development. WWGS-AIT is looking for a Data Scientist to design and build seller-facing AI agents that turn our AI-ready data foundation into intelligent, conversational experiences for Amazon's global sellers. You will own the intelligence layer of these agents end-to-end, from modeling and retrieval to evaluation and launch, working alongside applied scientists, data engineers, and the Seller Assistant platform team to put trustworthy AI directly into sellers' hands. Key job responsibilities - Design, build, and iterate seller-facing AI agents (LLM-powered) that help Chinese sellers grow globally, reasoning over WWGS-AIT's AI-ready data foundation and knowledge base. - Develop the intelligence layer of agents: retrieval-augmented generation (RAG) over our knowledge management system, tool-use / function-calling orchestration, prompt engineering, and model fine-tuning or adaptation where needed. - Ground agent responses in standardized metrics and unified seller profiles to guarantee consistency and accuracy across agents; design and enforce guardrails that prevent hallucination and protect sensitive, compliance-restricted data. - Build rigorous evaluation frameworks (golden datasets, offline evaluation, and online experimentation) to measure and continuously improve agent quality, safety, and seller impact. - Develop seller-intelligence models (segmentation, entity resolution / One-ID, ranking and recommendation) that power personalized agent experiences. - Partner with WWGS Tech and the Seller Assistant platform team to productionize agents and tools (e.g., via MCP), defining the model and intelligence contract while engineering operates the runtime. - Collaborate with business, product, and cross-functional partners to translate seller pain points into agent capabilities and measurable business outcomes. - Stay current with advances in GenAI and agentic systems, and bring applied research into production.