Card-Imbens 16x9.jpg
David Card (left), an Amazon Scholar, a professor of economics at the University of California, Berkeley, and the outgoing president of the AEA, and Guido Imbens (right), an academic research consultant at Amazon and a professor at the Stanford Graduate School of Business.

A conversation with economics Nobelists

Amazon Scholar David Card and academic research consultant Guido Imbens on the past and future of empirical economics.

The annual meeting of the American Economic Association (AEA) took place Jan. 7 - 9, and as it approached, Amazon Science had the chance to interview two of the three recipients of the 2021 Nobel Prize in economics — who also happen to be Amazon-affiliated economists.

David Card, an Amazon Scholar, a professor of economics at the University of California, Berkeley, and the outgoing president of the AEA, won half the prize “for his empirical contributions to labor economics”.

Guido Imbens, an academic research consultant at Amazon and a professor at the Stanford Graduate School of Business, shared the other half of the prize with MIT’s Josh Angrist for “methodological contributions to the analysis of causal relationships”.

Amazon Science: The empirical approach to economics has been recognized by the Nobel Prize committee several times in the last few years, but it wasn't always as popular as it is today. I'm curious how you both first became interested in empirical approaches to economics.

David Card: The heroes of economics for many, many decades were the theorists, and in the postwar era especially, there was a recognition that economic modeling was underdeveloped — the math was underdeveloped — and there was a need to formalize things and understand better what the models really delivered.

People started to realize that we had the data to better look at real labor market phenomena and possibly make economics something different than just a kind of a branch of philosophy.
David Card

That need really proceeded through the ’60s, and Arrow and Debreu were these famous mathematical economists who developed some very elegant theoretical models of how the market works in an idealized economy.

What happened in my time was people started to realize that we had the data to better look at real labor market phenomena and possibly make economics something different than just a kind of a branch of philosophy. Arrow-Debreu is basically mathematical philosophy.

Guido Imbens: I came from a very different tradition. I grew up in the Netherlands, and there was a strong tradition of econometrics started by people like Tinbergen. Tinbergen had been very broad — he did econometrics, but he also did empirical work and was very heavily involved in policy analysis. But over time, the program he had started was becoming much more focused on technical econometrics.

So as an undergraduate, we didn't really do any empirical work. We really just did a lot of mathematical statistics and some operations research and some economic theory. My thesis was a theoretical econometrics study.

When I presented that at Harvard, Josh Angrist wasn't really all that impressed with it, and he actually opposed the department hiring me there because he thought the paper was boring. And he was probably right! But luckily, the more senior people there at the time thought I was at least somewhat promising. And so I got hired at Harvard. But then it was really Josh and Larry Katz, one of the labor economists there, who got me interested in going to the labor seminar and got me exposed to the modern empirical work.

The context Josh and I started talking in really was this paper that I think came up in all three of the Nobel lectures, this paper by Ed Leamer, “Let's Take the Con Out of Econometrics”, where Leamer says, “Hardly anyone takes data analysis seriously. Or perhaps more accurately, hardly anyone takes anyone else’s data analysis seriously.”

And I think Leamer was right: people did these very elaborate things, and it was all showing off complicated technical things, but it wasn't really very credible. In fact, Leamer presented a lecture based on that work at Harvard. And I remember Josh getting up at some point and saying, “Well, you talk about all this old stuff, but look at the work Card does. Look at the work Krueger does. Look at the work I do. It's very different.”

And that felt right to me. It felt that the work was qualitatively very different from the work that Ed Leamer was describing and that he was complaining about.

AS: So that's when you first became aware of Professor Card’s work. Professor Card, when did you first become aware of Professor Imbens’s work?

Card: One of his early papers was pretty interesting. He was trying to combine data from micro survey evidence with benchmark numbers that you would get from a population, and it's actually a version of a kind of a problem that arises at Amazon all the time, which is, we've got noisy estimates of something, and we've got probably reliable estimates of some other aggregates, and there's often ways to try and combine those. I saw that and I thought that was very interesting.

Then there’s the problem that Josh and Guido worked on that was most impactful and that was cited by the Nobel Prize committee. I had worked on an experiment, a real experiment [as opposed to a natural experiment], in welfare analysis in Canada, and it was providing an economic incentive to try and get single mothers off of welfare and into work. And we noticed that the group of mothers who complied or followed on with the experiment was reasonable size, but it wasn't 100%.

We did some analysis of it trying to characterize them. Around the same time, I became aware of Imbens’s and Angrist’s paper, which basically formalized that a lot better and described what exactly was going on with this group. That framework just instantly took off, and everyone within a few years was thinking about problems that way.

This morning I was talking to another Amazon person about a problem. It was a difference analysis. I was saying we should try and characterize the compliers for this difference intervention. So it's exactly this problem.

The Nobel committee’s press release for Card, Imbens, and Angrist’s prize announcement emphasizes their use of natural experiments, which it defines as “situations in which chance events or policy changes result in groups of people being treated differently, in a way that resembles clinical trials in medicine.” A seminal instance of this was Card’s 1993 paper with his Princeton colleague Alan Krueger, which compared fast-food restaurants in two demographically similar communities on either side of the New Jersey-Pennsylvania border, one of which had recently seen a minimum-wage hike and one of which hadn’t.

AS: In the early days, there was skepticism about the empirical approach to economics. So every time you selected a new research project, you weren't just trying to answer an economics problem; you were also, in a sense, establishing the credibility of the approach. How did you select problems then? Was there a structure that you recognized as possibly lending itself to natural experiment?

Card: I think that the natural-experiment thing — there was really a brief period where that was novel, to tell you the truth. Maybe 1989 to 1992 or 3. I did this paper on the Mariel boatlift, which was cited by the committee. But to tell you the truth, that was a very modest paper. I never presented it anywhere, and it's in a very modest journal. So I never thought of that paper as going anywhere [laughs].

What happened was, it became more and more well understood that in order to make a claim of causality even from a natural-experiment setting, you had to have a fair amount of information from before the experiment took place to validate or verify that the group that you were calling the treatment group and the group that you were calling the control group actually were behaving the same.

That was a weakness of the project that Alan Krueger and I did. We had restaurants in New Jersey and Pennsylvania. We knew the minimum wage was going to increase — or we thought we knew that; it wasn't entirely clear at the time — but we surveyed the restaurants before, and then the minimum wage went up, and we surveyed them after, and that was good.

But we didn't really have multiple surveys from before to show that in the absence of the minimum wage, New Jersey and Pennsylvania restaurants had tracked each other for a long time. And these days, that's better understood. At Amazon for instance, people are doing intervention analyses of this type. They would normally look at what they call pre-trend analysis, make sure that the treatment group and the control group are trending the same beforehand.

I think there are 1,000 questions in economics that have been open forever. Sometimes new datasets come along. That's been happening a lot in labor economics: huge administrative datasets have become available, richer and richer, and now we're getting datasets that are created by these tech firms. So my usual thing is, I think, that's a dataset that maybe we can answer this old question on. That’s more my approach.

That's why being at Amazon has been great .... A lot of people have substantive questions they're trying to analyze with data, and they're kind of stuck in places, so there's a need for new methodologies.
Guido Imbens

Imbens: I come from a slightly different perspective. Most of my work has come from listening to people like David and Josh and seeing what type of problems they're working on, what type of methods they're using, and seeing if there's something to be added there — if there’s some way of improving the methods or places where maybe they're stuck, but listening to the people actually doing the empirical work rather than starting with the substantive questions.

That's why being at Amazon has been great, from my perspective. A lot of people have substantive questions they're trying to analyze with data, and they're kind of stuck in places, so there's a need for new methodologies. It's been a very fertile environment for me to come up with new research.

AS: Methodologically, what are some of the outstanding questions that interest you both?

Imbens: Well, one of the things is experimental design in complex environments. A lot of the experimental designs we’re using at the moment still come fairly directly from biomedical settings. We have a population, we randomize them into a treatment group and a control group, and then we compare outcomes for the two groups.

But in a lot of the settings we’re interested in at Amazon, there are very complex interactions between the units and their experiences, and dealing with that is very challenging. There are lots of special cases where we know somewhat what to do, but there are lots of cases where we don't know exactly what to do, and we need to do more complex experiments to get the answers to the questions we're interested in.

Double randomization — original color scheme.jpeg
An example of what Imbens calls “experimental design in complex environments”. In this illustration, each of five viewers is shown promotions for eight different Prime Video shows. Some of those promotions contain extra information, indicated in the image by star ratings (the “treatment”). This design helps determine whether the treatment affects viewing habits (the viewer experiment) but also helps identify spillover effects, in which participation in the viewer experiment influences the viewer’s behavior in other contexts.

The second thing is, we do a lot of these experiments, but often the experiments are relatively small. They’re small in duration, and they’re small in size relative to the overall population. You know, it goes back to the paper we mentioned before, combining this observational-study data with experimental data. That raises a lot of interesting methodological challenges that I spend a lot of time thinking about these days.

AS: I wondered if in the same way that in that early paper you were looking at survey data and population data, there's a way that natural experiments and economic field experiments can reinforce each other or give you a more reliable signal than you can get from either alone.

Card: There's one thing that people do; I've done a few of these myself. It's called meta analysis. It's a technique where you take results from different studies and try and put them into a statistical model. In a way it's comparable to work Guido has done at Amazon, where you take a series of actual experiments, A/B experiments done in Weblab, and basically combine them and say, “Okay, these aren't exactly the same products and the same conditions, but there's enough comparability that maybe I can build a model and use the information from the whole set to help inform what we're learning from any given one.”

And you can do that in studies in economics. For example, I’ve done one on training programs. There are many of these training programs. Each of them — exactly as Guido was saying — is often quite small. And there are weird conditions: sometimes it's only young males or young females that are in the experiment, or they don't have very long follow-up, or sometimes the labor market is really strong, and other times it's really weak. So you can try and build a model of the outcome you get from any given study and then try and see if there are any systematic patterns there.

Imbens: We do all these experiments, but often we kind of do them once, and then we put them aside. There's a lot of information over the years built up in all these experiments we've done, and finding more of these meta-analysis-type ways of combining them and exploiting all the information we have collected there — I think it's a very promising way to go.

AS: How can empirical methods complement theoretical approaches — model building of the kind that, in some sense, the early empirical research was reacting against?

Card: Normally, if you're building a model, there are a few key parameters, like you need to get some kind of an elasticity of what a customer will do if faced with a higher price or if offered a shorter, faster delivery speed versus slower delivery speed. And if you have those elasticities, then you can start building up a model.

If you have even a fairly complicated dynamic model, normally there's a relatively small number of these parameters, and the value of the model is to take this set of parameters and try and tell a bit richer story — not just how the customer responds to an offer of a faster delivery today but how that affects their future purchases and whether they come back and buy other products or whatever. But you need credible estimates of those elasticities. It's not helpful to build a model and then just pull numbers out of the air [laughs]. And that's why A/B experiments are so important at Amazon.

AS: I asked about outstanding methodological questions that you're interested in, but how about economic questions more broadly that you think could really benefit from an empirical approach?

Card: In my field [labor economics], we've begun to realize that different firms are setting different wages for the same kinds of workers. And we're starting to think about two issues related to that. One is, how do workers choose between jobs? Do they know about all the jobs out there? Do they just find out about some of the jobs? We're trying to figure out exactly why it's okay in the labor market for there to be multiple wages for a certain class of workers. Why don't all the workers immediately try to go to one job? This seems to be a very important phenomenon.

And on the other side of that, how do employers think about it? What are the benefits to employers of a higher wage or lower wage? Is it just the recruiting, or is it retention, or is it productivity? Is it longer-term goals? That's front and center in the research that I do outside of Amazon.

AS: I was curious if there were any cases where a problem presented itself, and at first you didn't think there was any way to get an empirical handle on it, and then you figured out that there was.

We're supposed to be social scientists who are trying to see what people are doing and the problems they confront and trying to analyze them. ... That's different than this old-fashioned Adam Smith view of the economy as a perfectly functioning tool that we're just supposed to admire.
David Card

Card: I saw a really interesting paper that was done by a PhD student who was visiting my center at Berkeley. In European football, there are a lot of non-white players, and fan racism is pretty pervasive. This guy noticed that during COVID, they played a lot of games with no fans. So he was able to compare the performance of the non-white and white players in the pre-COVID era and the COVID era, with and without fans, and showed that the non-white players did a little bit better. That's the kind of question where you’re saying, How are we ever going to study that? But if you're thinking and looking around, there's always some angle that might be useful.

Imbens: That's a very clever idea. I agree with David. If you just pay attention, there are a lot of things happening that allow you to answer important questions. Maybe fan insults in sports itself isn't that big a deal, but clearly, racism in the labor market and having people treated differently is a big problem. And here you get a very clear handle on an aspect of it. And once you show it's a problem there, it's very likely that it shows up in arguably substantively much more important settings where it's really hard to study.

In the Netherlands for a long time, they had a limit on the number of students who could go to medical school. And it wasn't decided by the medical schools themselves; they couldn't choose whom to admit. It was partly based on a lottery. At some point, someone used that to figure out how much access to medical school is actually worth. So essentially, you have two people who are both qualified to go to medical school; one gets lucky in the lottery; one doesn't. And it turns out you're giving the person who wins the lottery basically a lot of money. Obviously, in many professions we can't just randomly assign people to different types of jobs. But here you get a handle on the value of rationing that type of education.

Card: I think that's really important. You know, we're supposed to be social scientists who are trying to see what people are doing and the problems they confront and trying to analyze them. In a way, that's different than this sort of old-fashioned Adam Smith view of the economy as a perfectly functioning tool that we're just supposed to admire. That is a difference, I think.

Research areas

Related content

US, NY, New York
Fauna Robotics is building capable, safe, and delightful robots for everyday life, and voice is one of the most natural ways people will interact with them. Cloud speech and language models are good and getting better, but they can only work with the audio they receive, and a robot is a hard place to listen. Its microphones sit beside motors, fans, and moving joints. It speaks through its own loudspeaker while people talk over it. It moves, turns, and shares a room with several people at once. We are hiring a Principal Audio Scientist to be Fauna's technical authority on how our robots hear. You will design, prototype, and ship the hardest algorithms in the robot's audio system. You will set the audio architecture that other engineers build on, shape hardware decisions across robot generations, mentor the engineers and scientists working on audio, and be the person teams come to when the robot can't hear. Key job responsibilities - Set the long-range science roadmap and technical architecture for the robot's audio system, and serve as Fauna's primary technical authority on robot hearing - Design and implement suppression of the robot's own noise from motors, fans, and moving joints - Design and implement echo cancellation for the robot's own voice, so people can interrupt it naturally - Develop multi-microphone processing that holds up as the robot and the people around it move, including locating who is speaking so the robot can turn toward them - Make on-robot listening decisions robust to internal and external noise sources: wake word, voice activity, and whether speech is directed at the robot - Drive microphone and speaker placement, enclosure acoustics, and vibration isolation decisions with mechanical, electrical, and industrial design, backed by your own measurements - Design the robot-specific data collection and evaluation methods to validate the performance of our audio design - Present audio science and its tradeoffs to senior leadership and partner teams - Mentor scientists and engineers, raising the scientific bar for audio across the organization through design reviews, code reviews, and hiring
IN, KA, Bengaluru
Interested to build the next generation Financial systems that can handle billions of dollars in transactions? Interested to build highly scalable next generation systems that could utilize Amazon Cloud? Massive data volume + complex business rules in a highly distributed and service oriented architecture, a world class information collection and delivery challenge. Our challenge is to deliver the software systems which accurately capture, process, and report on the huge volume of financial transactions that are generated each day as millions of customers make purchases, as thousands of Vendors and Partners are paid, as inventory moves in and out of warehouses, as commissions are calculated, and as taxes are collected in hundreds of jurisdictions worldwide. Key job responsibilities • Understand the business and discover actionable insights from large volumes of data through application of machine learning, statistics or causal inference. • Analyse and extract relevant information from large amounts of Amazon’s historical transactions data to help automate and optimize key processes • Research, develop and implement novel machine learning and statistical approaches for anomaly, theft, fraud, abusive and wasteful transactions detection. • Use machine learning and analytical techniques to create scalable solutions for business problems. • Identify new areas where machine learning can be applied for solving business problems. • Partner with developers and business teams to put your models in production. • Mentor other scientists and engineers in the use of ML techniques. A day in the life • Understand the business and discover actionable insights from large volumes of data through application of machine learning, statistics or causal inference. • Analyse and extract relevant information from large amounts of Amazon’s historical transactions data to help automate and optimize key processes • Research, develop and implement novel machine learning and statistical approaches for anomaly, theft, fraud, abusive and wasteful transactions detection. • Use machine learning and analytical techniques to create scalable solutions for business problems. • Identify new areas where machine learning can be applied for solving business problems. • Partner with developers and business teams to put your models in production. • Mentor other scientists and engineers in the use of ML techniques. About the team The FinAuto TFAW(theft, fraud, abuse, waste) team is part of FGBS Org and focuses on building applications utilizing machine learning models to identify and prevent theft, fraud, abusive and wasteful(TFAW) financial transactions across Amazon. Our mission is to prevent every single TFAW transaction. As a Machine Learning Scientist in the team, you will be driving the TFAW Sciences roadmap, conduct research to develop state-of-the-art solutions through a combination of data mining, statistical and machine learning techniques, and coordinate with Engineering team to put these models into production. You will need to collaborate effectively with internal stakeholders, cross-functional teams to solve problems, create operational efficiencies, and deliver successfully against high organizational standards.
US, NY, New York
We are seeking a Robotics/AI Motor Control Scientist to develop cutting-edge machine learning algorithms for motor control systems in robots. In this role, you will focus on creating and optimizing intelligent motor control strategies to enable robots to perform complex, whole-body tasks. Your contributions will be essential in advancing robotics by enabling fluid, reliable, and safe interactions between robots and their environments. Key job responsibilities - Develop controllers that leverage reinforcement learning, imitation learning, or other advanced AI techniques to achieve natural, robust, and adaptive motor behaviors - Collaborate with multi-disciplinary teams to integrate motor control systems with robotic hardware, ensuring alignment with real-world constraints such as actuator dynamics and energy efficiency - Use simulation and real-world testing to refine and validate control algorithms - Stay updated on advancements in robotics, AI, and control systems to apply advanced techniques to robotic motion challenges - Lead technical projects from conception through production deployment - Mentor junior scientists and engineers - Bridge research initiatives with practical engineering implementation About the team Fauna Robotics, an Amazon company, is building capable, safe, and genuinely delightful robots for everyday life. Our goal is simple: make robots people actually want to live and interact with in everyday human spaces. We believe that future won’t arrive until building for robotics becomes far more accessible. Today, too much effort is spent reinventing the fundamentals. We’re changing that by developing tightly integrated hardware and software systems that make it faster, safer, and more intuitive to create real-world robotic products. Our work spans the full stack: mechanical design, control systems, dynamic modeling, and intelligent software. The focus is not just functionality, but experience. We’re building robots that feel responsive, expressive, and genuinely useful. At Fauna, you’ll work at the frontier of this space, helping define how robots move, manipulate, and interact with people in natural environments. It’s an opportunity to solve hard problems across hardware and software with a team focused on making robotics accessible and joyful to build. If you care about making robotics real for everyone and building systems that are as delightful as they are capable, we’re interested in hearing from you. an opportunity to solve hard problems across hardware and software with a team focused on making robotics accessible and joyful to build. If you care about making robotics real for everyone and building systems that are as delightful as they are capable, we’re interested in hearing from you.
US, CA, Sunnyvale
Amazon's Artificial General Intelligence (AGI) organization is seeking an Applied Scientist III to advance the science of Responsible AI evaluation for large language models and generative AI. In this role, you will lead the design and development of rigorous evaluation methods, benchmarks, and metrics that measure the safety, fairness, robustness, and trustworthiness of frontier models. You will work with large-scale datasets, modern deep learning frameworks, and world-class scientists and engineers to turn research into evaluation systems that shape model launch decisions at Amazon scale. Key job responsibilities - Lead the design and implementation of evaluation frameworks, benchmarks, and metrics for responsible AI, including safety, fairness, robustness, and harmful content. - Build scalable automated evaluation pipelines for large language models, including model-based and human-in-the-loop evaluation. - Partner with pretraining, post-training, and product teams to translate evaluation results into model improvements and launch decisions. - Conduct rigorous experimentation and statistical analysis, and publish research at top venues. - Mentor junior scientists and help raise the scientific bar of the team. - Champion responsible AI practices across the model development lifecycle. About the team The AGI Responsible AI (RAI) team builds the science and systems that make Amazon's large language models safe, fair, and trustworthy. We work on problems spanning safety evaluation, content moderation, watermarking, bias mitigation, and alignment. Our team values scientific rigor, customer obsession, and rapid iteration, and we collaborate closely with pretraining, post-training, and product teams across AGI.
CH, Zurich
RIVR, an Amazon company, is building Physical AI by deploying autonomous robots for real-world doorstep delivery. Operating daily in diverse urban environments, RIVR's robots continuously learn from and navigate the millions of scenarios encountered during deliveries. By owning the full stack from software to hardware, RIVR is purpose-built for safety, reliability, and the customer from day one. Reinforcement learning is transforming our robotic intelligence, enabling autonomous behavior without human guidance. We are seeking a Senior AI Engineer with deep expertise in reinforcement learning and deep learning, including supervised and self-supervised learning with a focus on dexterous manipulation. Your role will involve leveraging both simulated and real-world data to address practical challenges in dynamic grasping, contact-rich manipulation, and object interaction. If you are passionate about advancing AI and developing innovative solutions, join us in shaping the future of intelligent robotics. Key job responsibilities Develop cutting-edge reinforcement learning algorithms to enable robust, contact-rich dexterous manipulation, translating vision, depth, tactile, and proprioceptive sensor input into precise end-effector and joint-level motor commands. Design, test, and refine algorithms to solve complex real-world manipulation challenges, such as handling diverse package form factors, dynamic hand-offs, and operating door handles or latches. Collaborate with the foundation model team to innovate methods that leverage both simulated and real-world data.
US, WA, Seattle
Do you want to join an innovative team of scientists who use machine learning and statistical techniques to help Amazon provide the best customer experience by preventing eCommerce fraud? Are you excited by the prospect of analyzing and modeling terabytes of data and creating state-of-the-art algorithms to solve real world problems? Do you like to own end-to-end business problems/metrics and directly impact the profitability of the company? Do you enjoy collaborating in a diverse team environment? If yes, then you may be a great fit to join the Amazon Selling Partner Trust & Store Integrity Science Team. We are looking for a talented scientist who is passionate to build advanced machine learning systems that help manage the safety of millions of transactions every day and scale up our operation with automation. Key job responsibilities Innovate with the latest GenAI/LLM/VLM technology to build highly automated solutions for efficient fraud detection, risk evaluation and automated operations Design, develop and deploy end-to-end advance machine learning solutions with vision anf GenAi technologies in the Amazon production environment to create impactful business value Learn, explore and experiment with the latest machine learning advancements to create the best customer experience A day in the life You will be working within a dynamic, diverse, and supportive group of scientists who share your passion for innovation and excellence. You'll be working closely with business partners and engineering teams to create end-to-end scalable machine learning solutions that address real-world problems. You will build scalable, efficient, and automated processes for large-scale data analyses, model development, model validation, and model implementation. You will also be providing clear and compelling reports for your solutions and contributing to the ongoing innovation and knowledge-sharing that are central to the team's success.
US, NY, New York
We are seeking an Applied Scientist to contribute to research and development of novel security validation and monitoring techniques for AI systems at scale. You will own and contribute to four critical work-streams: 1. Real-Time Agent Monitoring Design and implement scientific approaches for continuous behavioral analysis of AI agents in production—detecting anomalous actions, prompt injection exploitation, and policy violations in real time. 2. Protection & Automated Remediation Invent and deliver novel protection technologies and automated remediation techniques building on research in security, cryptography, privacy, automated reasoning, and others domains, to enable safe and secure agentic AI models and AI applications. 3. AI Application and Capabilities Validation Invent and deliver scalable methodologies for security testing of AI applications and AI capabilities (e.g. MCP, skills), including adversarial robustness evaluation, safety guardrail bypass detection, tool-use authorization boundaries, and trust boundary verification. 4. AI Asset Discovery & Inventory Research and build scalable techniques to automatically discover, identify, and catalog all AI-enabled applications, services, and capabilities across the company—maintaining a comprehensive, continuously updated database of AI assets. Key job responsibilities Invent • Identify and frame new research challenges in AI security where problems are ill-defined and require novel scientific paradigms at the product level. • Contribute to the team's scientific agenda for agent monitoring, protection, remediation, validation research, and AI asset discovery. • Publish research results at peer-reviewed internal and external venues (e.g., USENIX Security, ACM CCS, IEEE S&P, NeurIPS, ICML security workshops, ICSE, PETS) when appropriate. • Articulate key scientific challenges of current and future AI security threats and deliver novel research to address them. • Design, implementation, and successful delivery of scientifically complex security solutions into production—both brand new systems and evolutions of existing ones. • Write significant portions of critical-path code for detection models, validation engines, protection technologies, and asset discovery / classification systems. • Assess and select appropriate technologies (e.g., data protection, private inference, graph-based anomaly detection, NLP-based service classification, code/traffic analysis for AI fingerprinting) for production systems. • Use best practices in scientific methodology and software engineering across the team; provide insightful peer reviews of code, design, and architecture artifacts. • Deliver solutions that are inventive, maintainable, scalable, and extensible. Influence • Autonomously drive discussions with security engineers, software engineers, product managers, and scientist peers across multiple teams. • Build consensus on larger cross-team security initiatives and factor complex efforts into independent workstreams. • Identify and resolve endemic problems, including areas where current security tooling limits innovation of partner teams. • Contribute to the broader internal and external scientific communities as a subject matter expert in AI security. About the team Diverse Experiences Amazon Security values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying. Why Amazon Security? At Amazon, security is central to maintaining customer trust and delivering delightful customer experiences. Our organization is responsible for creating and maintaining a high bar for security across all of Amazon’s products and services. We offer talented security professionals the chance to accelerate their careers with opportunities to build experience in a wide variety of areas including cloud, devices, retail, entertainment, healthcare, operations, and physical stores. Inclusive Team Culture In Amazon Security, it’s in our nature to learn and be curious. Ongoing DEI events and learning experiences inspire us to continue learning and to embrace our uniqueness. Addressing the toughest security challenges requires that we seek out and celebrate a diversity of ideas, perspectives, and voices. Training & Career Growth We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, training, and other career-advancing resources here to help you develop into a better-rounded professional. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why flexible work hours and arrangements are part of our culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve.
IN, KA, Bangalore
Amazon’s Last Mile Team is looking for a passionate individual with strong optimization and analytical skills to join its Last Mile Science team in the endeavor of designing and improving the most complex planning of delivery network in the world. Last Mile builds global solutions that enable Amazon to attract an elastic supply of drivers, companies, and assets needed to deliver Amazon's and other shippers' volumes at the lowest cost and with the best customer delivery experience. Last Mile Science team owns the core decision models in the space of jurisdiction planning, delivery channel and modes network design, capacity planning for on the road and at delivery stations, routing inputs estimation and optimization. Our research has direct impact on customer experience, driver and station associate experience, Delivery Service Partner (DSP)’s success and the sustainable growth of Amazon. Optimizing the last mile delivery requires deep understanding of transportation, supply chain management, pricing strategies and forecasting. Only through innovative and strategic thinking, we will make the right capital investments in technology, assets and infrastructures that allows for long-term success. Our team members have an opportunity to be on the forefront of supply chain thought leadership by working on some of the most difficult problems in the industry with some of the best product managers, scientists, and software engineers in the industry. Key job responsibilities Candidates will be responsible for developing solutions to better manage and optimize delivery capacity in the last mile network. The successful candidate should have solid research experience in one or more technical areas of Operations Research or Machine Learning. These positions will focus on identifying and analyzing opportunities to improve existing algorithms and also on optimizing the system policies across the management of external delivery service providers and internal planning strategies. They require superior logical thinkers who are able to quickly approach large ambiguous problems, turn high-level business requirements into mathematical models, identify the right solution approach, and contribute to the software development for production systems. To support their proposals, candidates should be able to independently mine and analyze data, and be able to use any necessary programming and statistical analysis software to do so. Successful candidates must thrive in fast-paced environments, which encourage collaborative and creative problem solving, be able to measure and estimate risks, constructively critique peer research, and align research focuses with the Amazon's strategic needs. As a senior scientist, you will also help coach/mentor junior scientists in the team.
US, NY, New York
We are seeking an Applied Scientist to lead the development of evaluation frameworks and data collection protocols for robotic capabilities. In this role, you will focus on designing how we measure, stress-test, and improve robot behavior across a wide range of real-world tasks. Your work will play a critical role in shaping how policies are validated and how high-quality datasets are generated to accelerate system performance. You will operate at the intersection of robotics, machine learning, and human-in-the-loop systems, building the infrastructure and methodologies that connect teleoperation, evaluation, and learning. This includes developing evaluation policies, defining task structures, and contributing to operator-facing interfaces that enable scalable and reliable data collection. The ideal candidate is highly experimental, systems-oriented, and comfortable working across software, robotics, and data pipelines, with a strong focus on turning ambiguous capability goals into measurable and actionable evaluation systems. Key job responsibilities - Design and implement evaluation frameworks to measure robot capabilities across structured tasks, edge cases, and real-world scenarios - Develop task definitions, success criteria, and benchmarking methodologies that enable consistent and reproducible evaluation of policies - Create and refine data collection protocols that generate high-quality, task-relevant datasets aligned with model development needs - Build and iterate on teleoperation workflows and operator interfaces to support efficient, reliable, and scalable data collection - Analyze evaluation results and collected data to identify performance gaps, failure modes, and opportunities for targeted data collection - Collaborate with engineering teams to integrate evaluation tooling, logging systems, and data pipelines into the broader robotics stack - Stay current with advances in robotics, evaluation methodologies, and human-in-the-loop learning to continuously improve internal approaches - Lead technical projects from conception through production deployment - Mentor junior scientists and engineers
US, NY, New York
We are seeking an Applied Scientist to lead the development of evaluation frameworks and data collection protocols for robotic capabilities. In this role, you will focus on designing how we measure, stress-test, and improve robot behavior across a wide range of real-world tasks. Your work will play a critical role in shaping how policies are validated and how high-quality datasets are generated to accelerate system performance. You will operate at the intersection of robotics, machine learning, and human-in-the-loop systems, building the infrastructure and methodologies that connect teleoperation, evaluation, and learning. This includes developing evaluation policies, defining task structures, and contributing to operator-facing interfaces that enable scalable and reliable data collection. The ideal candidate is highly experimental, systems-oriented, and comfortable working across software, robotics, and data pipelines, with a strong focus on turning ambiguous capability goals into measurable and actionable evaluation systems. Key job responsibilities - Design and implement evaluation frameworks to measure robot capabilities across structured tasks, edge cases, and real-world scenarios - Develop task definitions, success criteria, and benchmarking methodologies that enable consistent and reproducible evaluation of policies - Create and refine data collection protocols that generate high-quality, task-relevant datasets aligned with model development needs - Build and iterate on teleoperation workflows and operator interfaces to support efficient, reliable, and scalable data collection - Analyze evaluation results and collected data to identify performance gaps, failure modes, and opportunities for targeted data collection - Collaborate with engineering teams to integrate evaluation tooling, logging systems, and data pipelines into the broader robotics stack - Stay current with advances in robotics, evaluation methodologies, and human-in-the-loop learning to continuously improve internal approaches - Lead technical projects from conception through production deployment - Mentor junior scientists and engineers