Card-Imbens 16x9.jpg
David Card (left), an Amazon Scholar, a professor of economics at the University of California, Berkeley, and the outgoing president of the AEA, and Guido Imbens (right), an academic research consultant at Amazon and a professor at the Stanford Graduate School of Business.

A conversation with economics Nobelists

Amazon Scholar David Card and academic research consultant Guido Imbens on the past and future of empirical economics.

The annual meeting of the American Economic Association (AEA) took place Jan. 7 - 9, and as it approached, Amazon Science had the chance to interview two of the three recipients of the 2021 Nobel Prize in economics — who also happen to be Amazon-affiliated economists.

David Card, an Amazon Scholar, a professor of economics at the University of California, Berkeley, and the outgoing president of the AEA, won half the prize “for his empirical contributions to labor economics”.

Guido Imbens, an academic research consultant at Amazon and a professor at the Stanford Graduate School of Business, shared the other half of the prize with MIT’s Josh Angrist for “methodological contributions to the analysis of causal relationships”.

Amazon Science: The empirical approach to economics has been recognized by the Nobel Prize committee several times in the last few years, but it wasn't always as popular as it is today. I'm curious how you both first became interested in empirical approaches to economics.

David Card: The heroes of economics for many, many decades were the theorists, and in the postwar era especially, there was a recognition that economic modeling was underdeveloped — the math was underdeveloped — and there was a need to formalize things and understand better what the models really delivered.

People started to realize that we had the data to better look at real labor market phenomena and possibly make economics something different than just a kind of a branch of philosophy.
David Card

That need really proceeded through the ’60s, and Arrow and Debreu were these famous mathematical economists who developed some very elegant theoretical models of how the market works in an idealized economy.

What happened in my time was people started to realize that we had the data to better look at real labor market phenomena and possibly make economics something different than just a kind of a branch of philosophy. Arrow-Debreu is basically mathematical philosophy.

Guido Imbens: I came from a very different tradition. I grew up in the Netherlands, and there was a strong tradition of econometrics started by people like Tinbergen. Tinbergen had been very broad — he did econometrics, but he also did empirical work and was very heavily involved in policy analysis. But over time, the program he had started was becoming much more focused on technical econometrics.

So as an undergraduate, we didn't really do any empirical work. We really just did a lot of mathematical statistics and some operations research and some economic theory. My thesis was a theoretical econometrics study.

When I presented that at Harvard, Josh Angrist wasn't really all that impressed with it, and he actually opposed the department hiring me there because he thought the paper was boring. And he was probably right! But luckily, the more senior people there at the time thought I was at least somewhat promising. And so I got hired at Harvard. But then it was really Josh and Larry Katz, one of the labor economists there, who got me interested in going to the labor seminar and got me exposed to the modern empirical work.

The context Josh and I started talking in really was this paper that I think came up in all three of the Nobel lectures, this paper by Ed Leamer, “Let's Take the Con Out of Econometrics”, where Leamer says, “Hardly anyone takes data analysis seriously. Or perhaps more accurately, hardly anyone takes anyone else’s data analysis seriously.”

And I think Leamer was right: people did these very elaborate things, and it was all showing off complicated technical things, but it wasn't really very credible. In fact, Leamer presented a lecture based on that work at Harvard. And I remember Josh getting up at some point and saying, “Well, you talk about all this old stuff, but look at the work Card does. Look at the work Krueger does. Look at the work I do. It's very different.”

And that felt right to me. It felt that the work was qualitatively very different from the work that Ed Leamer was describing and that he was complaining about.

AS: So that's when you first became aware of Professor Card’s work. Professor Card, when did you first become aware of Professor Imbens’s work?

Card: One of his early papers was pretty interesting. He was trying to combine data from micro survey evidence with benchmark numbers that you would get from a population, and it's actually a version of a kind of a problem that arises at Amazon all the time, which is, we've got noisy estimates of something, and we've got probably reliable estimates of some other aggregates, and there's often ways to try and combine those. I saw that and I thought that was very interesting.

Then there’s the problem that Josh and Guido worked on that was most impactful and that was cited by the Nobel Prize committee. I had worked on an experiment, a real experiment [as opposed to a natural experiment], in welfare analysis in Canada, and it was providing an economic incentive to try and get single mothers off of welfare and into work. And we noticed that the group of mothers who complied or followed on with the experiment was reasonable size, but it wasn't 100%.

We did some analysis of it trying to characterize them. Around the same time, I became aware of Imbens’s and Angrist’s paper, which basically formalized that a lot better and described what exactly was going on with this group. That framework just instantly took off, and everyone within a few years was thinking about problems that way.

This morning I was talking to another Amazon person about a problem. It was a difference analysis. I was saying we should try and characterize the compliers for this difference intervention. So it's exactly this problem.

The Nobel committee’s press release for Card, Imbens, and Angrist’s prize announcement emphasizes their use of natural experiments, which it defines as “situations in which chance events or policy changes result in groups of people being treated differently, in a way that resembles clinical trials in medicine.” A seminal instance of this was Card’s 1993 paper with his Princeton colleague Alan Krueger, which compared fast-food restaurants in two demographically similar communities on either side of the New Jersey-Pennsylvania border, one of which had recently seen a minimum-wage hike and one of which hadn’t.

AS: In the early days, there was skepticism about the empirical approach to economics. So every time you selected a new research project, you weren't just trying to answer an economics problem; you were also, in a sense, establishing the credibility of the approach. How did you select problems then? Was there a structure that you recognized as possibly lending itself to natural experiment?

Card: I think that the natural-experiment thing — there was really a brief period where that was novel, to tell you the truth. Maybe 1989 to 1992 or 3. I did this paper on the Mariel boatlift, which was cited by the committee. But to tell you the truth, that was a very modest paper. I never presented it anywhere, and it's in a very modest journal. So I never thought of that paper as going anywhere [laughs].

What happened was, it became more and more well understood that in order to make a claim of causality even from a natural-experiment setting, you had to have a fair amount of information from before the experiment took place to validate or verify that the group that you were calling the treatment group and the group that you were calling the control group actually were behaving the same.

That was a weakness of the project that Alan Krueger and I did. We had restaurants in New Jersey and Pennsylvania. We knew the minimum wage was going to increase — or we thought we knew that; it wasn't entirely clear at the time — but we surveyed the restaurants before, and then the minimum wage went up, and we surveyed them after, and that was good.

But we didn't really have multiple surveys from before to show that in the absence of the minimum wage, New Jersey and Pennsylvania restaurants had tracked each other for a long time. And these days, that's better understood. At Amazon for instance, people are doing intervention analyses of this type. They would normally look at what they call pre-trend analysis, make sure that the treatment group and the control group are trending the same beforehand.

I think there are 1,000 questions in economics that have been open forever. Sometimes new datasets come along. That's been happening a lot in labor economics: huge administrative datasets have become available, richer and richer, and now we're getting datasets that are created by these tech firms. So my usual thing is, I think, that's a dataset that maybe we can answer this old question on. That’s more my approach.

That's why being at Amazon has been great .... A lot of people have substantive questions they're trying to analyze with data, and they're kind of stuck in places, so there's a need for new methodologies.
Guido Imbens

Imbens: I come from a slightly different perspective. Most of my work has come from listening to people like David and Josh and seeing what type of problems they're working on, what type of methods they're using, and seeing if there's something to be added there — if there’s some way of improving the methods or places where maybe they're stuck, but listening to the people actually doing the empirical work rather than starting with the substantive questions.

That's why being at Amazon has been great, from my perspective. A lot of people have substantive questions they're trying to analyze with data, and they're kind of stuck in places, so there's a need for new methodologies. It's been a very fertile environment for me to come up with new research.

AS: Methodologically, what are some of the outstanding questions that interest you both?

Imbens: Well, one of the things is experimental design in complex environments. A lot of the experimental designs we’re using at the moment still come fairly directly from biomedical settings. We have a population, we randomize them into a treatment group and a control group, and then we compare outcomes for the two groups.

But in a lot of the settings we’re interested in at Amazon, there are very complex interactions between the units and their experiences, and dealing with that is very challenging. There are lots of special cases where we know somewhat what to do, but there are lots of cases where we don't know exactly what to do, and we need to do more complex experiments to get the answers to the questions we're interested in.

Double randomization — original color scheme.jpeg
An example of what Imbens calls “experimental design in complex environments”. In this illustration, each of five viewers is shown promotions for eight different Prime Video shows. Some of those promotions contain extra information, indicated in the image by star ratings (the “treatment”). This design helps determine whether the treatment affects viewing habits (the viewer experiment) but also helps identify spillover effects, in which participation in the viewer experiment influences the viewer’s behavior in other contexts.

The second thing is, we do a lot of these experiments, but often the experiments are relatively small. They’re small in duration, and they’re small in size relative to the overall population. You know, it goes back to the paper we mentioned before, combining this observational-study data with experimental data. That raises a lot of interesting methodological challenges that I spend a lot of time thinking about these days.

AS: I wondered if in the same way that in that early paper you were looking at survey data and population data, there's a way that natural experiments and economic field experiments can reinforce each other or give you a more reliable signal than you can get from either alone.

Card: There's one thing that people do; I've done a few of these myself. It's called meta analysis. It's a technique where you take results from different studies and try and put them into a statistical model. In a way it's comparable to work Guido has done at Amazon, where you take a series of actual experiments, A/B experiments done in Weblab, and basically combine them and say, “Okay, these aren't exactly the same products and the same conditions, but there's enough comparability that maybe I can build a model and use the information from the whole set to help inform what we're learning from any given one.”

And you can do that in studies in economics. For example, I’ve done one on training programs. There are many of these training programs. Each of them — exactly as Guido was saying — is often quite small. And there are weird conditions: sometimes it's only young males or young females that are in the experiment, or they don't have very long follow-up, or sometimes the labor market is really strong, and other times it's really weak. So you can try and build a model of the outcome you get from any given study and then try and see if there are any systematic patterns there.

Imbens: We do all these experiments, but often we kind of do them once, and then we put them aside. There's a lot of information over the years built up in all these experiments we've done, and finding more of these meta-analysis-type ways of combining them and exploiting all the information we have collected there — I think it's a very promising way to go.

AS: How can empirical methods complement theoretical approaches — model building of the kind that, in some sense, the early empirical research was reacting against?

Card: Normally, if you're building a model, there are a few key parameters, like you need to get some kind of an elasticity of what a customer will do if faced with a higher price or if offered a shorter, faster delivery speed versus slower delivery speed. And if you have those elasticities, then you can start building up a model.

If you have even a fairly complicated dynamic model, normally there's a relatively small number of these parameters, and the value of the model is to take this set of parameters and try and tell a bit richer story — not just how the customer responds to an offer of a faster delivery today but how that affects their future purchases and whether they come back and buy other products or whatever. But you need credible estimates of those elasticities. It's not helpful to build a model and then just pull numbers out of the air [laughs]. And that's why A/B experiments are so important at Amazon.

AS: I asked about outstanding methodological questions that you're interested in, but how about economic questions more broadly that you think could really benefit from an empirical approach?

Card: In my field [labor economics], we've begun to realize that different firms are setting different wages for the same kinds of workers. And we're starting to think about two issues related to that. One is, how do workers choose between jobs? Do they know about all the jobs out there? Do they just find out about some of the jobs? We're trying to figure out exactly why it's okay in the labor market for there to be multiple wages for a certain class of workers. Why don't all the workers immediately try to go to one job? This seems to be a very important phenomenon.

And on the other side of that, how do employers think about it? What are the benefits to employers of a higher wage or lower wage? Is it just the recruiting, or is it retention, or is it productivity? Is it longer-term goals? That's front and center in the research that I do outside of Amazon.

AS: I was curious if there were any cases where a problem presented itself, and at first you didn't think there was any way to get an empirical handle on it, and then you figured out that there was.

We're supposed to be social scientists who are trying to see what people are doing and the problems they confront and trying to analyze them. ... That's different than this old-fashioned Adam Smith view of the economy as a perfectly functioning tool that we're just supposed to admire.
David Card

Card: I saw a really interesting paper that was done by a PhD student who was visiting my center at Berkeley. In European football, there are a lot of non-white players, and fan racism is pretty pervasive. This guy noticed that during COVID, they played a lot of games with no fans. So he was able to compare the performance of the non-white and white players in the pre-COVID era and the COVID era, with and without fans, and showed that the non-white players did a little bit better. That's the kind of question where you’re saying, How are we ever going to study that? But if you're thinking and looking around, there's always some angle that might be useful.

Imbens: That's a very clever idea. I agree with David. If you just pay attention, there are a lot of things happening that allow you to answer important questions. Maybe fan insults in sports itself isn't that big a deal, but clearly, racism in the labor market and having people treated differently is a big problem. And here you get a very clear handle on an aspect of it. And once you show it's a problem there, it's very likely that it shows up in arguably substantively much more important settings where it's really hard to study.

In the Netherlands for a long time, they had a limit on the number of students who could go to medical school. And it wasn't decided by the medical schools themselves; they couldn't choose whom to admit. It was partly based on a lottery. At some point, someone used that to figure out how much access to medical school is actually worth. So essentially, you have two people who are both qualified to go to medical school; one gets lucky in the lottery; one doesn't. And it turns out you're giving the person who wins the lottery basically a lot of money. Obviously, in many professions we can't just randomly assign people to different types of jobs. But here you get a handle on the value of rationing that type of education.

Card: I think that's really important. You know, we're supposed to be social scientists who are trying to see what people are doing and the problems they confront and trying to analyze them. In a way, that's different than this sort of old-fashioned Adam Smith view of the economy as a perfectly functioning tool that we're just supposed to admire. That is a difference, I think.

Research areas

Related content

US, CA, Pasadena
The Amazon Center for Quantum Computing in Pasadena, CA, is looking to hire an Applied Scientist specializing in Mixed-Signal Design. Working alongside other scientists and engineers, you will design and validate hardware performing the control and readout functions for AWS quantum processors. Candidates must have a solid background in mixed-signal design at the printed circuit board (PCB) level. Working effectively within a cross-functional team environment is critical. The ideal candidate will have demonstrated the capability to contribute to all phases of product life cycle development, from requirements gathering to verification. Diverse Experiences Amazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud. Inclusive Team Culture Here at Amazon, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness. Mentorship and Career Growth We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional. Key job responsibilities Our scientists and engineers collaborate across diverse teams and projects to offer state of the art, cost effective solutions for the control of Amazon quantum processor systems. You’ll bring a passion for innovation, collaboration, and mentoring to: Solve layered technical problems, often ones not encountered before, across our hardware stack. Develop requirements with key system stakeholders, including quantum device, test and measurement, and cryogenic hardware teams. Design, implement, test, deploy, and maintain innovative solutions that meet both strict performance and cost metrics. Research enabling control system technologies necessary for Amazon to produce commercially viable quantum computers.
IN, KA, Bengaluru
Interested to build the next generation Financial systems that can handle billions of dollars in transactions? Interested to build highly scalable next generation systems that could utilize Amazon Cloud? Massive data volume + complex business rules in a highly distributed and service oriented architecture, a world class information collection and delivery challenge. Our challenge is to deliver the software systems which accurately capture, process, and report on the huge volume of financial transactions that are generated each day as millions of customers make purchases, as thousands of Vendors and Partners are paid, as inventory moves in and out of warehouses, as commissions are calculated, and as taxes are collected in hundreds of jurisdictions worldwide. Key job responsibilities • Understand the business and discover actionable insights from large volumes of data through application of machine learning, statistics or causal inference. • Analyse and extract relevant information from large amounts of Amazon’s historical transactions data to help automate and optimize key processes • Research, develop and implement novel machine learning and statistical approaches for anomaly, theft, fraud, abusive and wasteful transactions detection. • Use machine learning and analytical techniques to create scalable solutions for business problems. • Identify new areas where machine learning can be applied for solving business problems. • Partner with developers and business teams to put your models in production. • Mentor other scientists and engineers in the use of ML techniques. A day in the life • Understand the business and discover actionable insights from large volumes of data through application of machine learning, statistics or causal inference. • Analyse and extract relevant information from large amounts of Amazon’s historical transactions data to help automate and optimize key processes • Research, develop and implement novel machine learning and statistical approaches for anomaly, theft, fraud, abusive and wasteful transactions detection. • Use machine learning and analytical techniques to create scalable solutions for business problems. • Identify new areas where machine learning can be applied for solving business problems. • Partner with developers and business teams to put your models in production. • Mentor other scientists and engineers in the use of ML techniques. About the team The FinAuto TFAW(theft, fraud, abuse, waste) team is part of FGBS Org and focuses on building applications utilizing machine learning models to identify and prevent theft, fraud, abusive and wasteful(TFAW) financial transactions across Amazon. Our mission is to prevent every single TFAW transaction. As a Machine Learning Scientist in the team, you will be driving the TFAW Sciences roadmap, conduct research to develop state-of-the-art solutions through a combination of data mining, statistical and machine learning techniques, and coordinate with Engineering team to put these models into production. You will need to collaborate effectively with internal stakeholders, cross-functional teams to solve problems, create operational efficiencies, and deliver successfully against high organizational standards.
IN, KA, Bengaluru
Interested to build the next generation Financial systems that can handle billions of dollars in transactions? Interested to build highly scalable next generation systems that could utilize Amazon Cloud? Massive data volume + complex business rules in a highly distributed and service oriented architecture, a world class information collection and delivery challenge. Our challenge is to deliver the software systems which accurately capture, process, and report on the huge volume of financial transactions that are generated each day as millions of customers make purchases, as thousands of Vendors and Partners are paid, as inventory moves in and out of warehouses, as commissions are calculated, and as taxes are collected in hundreds of jurisdictions worldwide. Key job responsibilities • Understand the business and discover actionable insights from large volumes of data through application of machine learning, statistics or causal inference. • Analyse and extract relevant information from large amounts of Amazon’s historical transactions data to help automate and optimize key processes • Research, develop and implement novel machine learning and statistical approaches for anomaly, theft, fraud, abusive and wasteful transactions detection. • Use machine learning and analytical techniques to create scalable solutions for business problems. • Identify new areas where machine learning can be applied for solving business problems. • Partner with developers and business teams to put your models in production. • Mentor other scientists and engineers in the use of ML techniques. A day in the life • Understand the business and discover actionable insights from large volumes of data through application of machine learning, statistics or causal inference. • Analyse and extract relevant information from large amounts of Amazon’s historical transactions data to help automate and optimize key processes • Research, develop and implement novel machine learning and statistical approaches for anomaly, theft, fraud, abusive and wasteful transactions detection. • Use machine learning and analytical techniques to create scalable solutions for business problems. • Identify new areas where machine learning can be applied for solving business problems. • Partner with developers and business teams to put your models in production. • Mentor other scientists and engineers in the use of ML techniques. About the team The FinAuto TFAW(theft, fraud, abuse, waste) team is part of FGBS Org and focuses on building applications utilizing machine learning models to identify and prevent theft, fraud, abusive and wasteful(TFAW) financial transactions across Amazon. Our mission is to prevent every single TFAW transaction. As a Machine Learning Scientist in the team, you will be driving the TFAW Sciences roadmap, conduct research to develop state-of-the-art solutions through a combination of data mining, statistical and machine learning techniques, and coordinate with Engineering team to put these models into production. You will need to collaborate effectively with internal stakeholders, cross-functional teams to solve problems, create operational efficiencies, and deliver successfully against high organizational standards.
IN, KA, Bengaluru
Amazon Health Services (One Medical) About Us: At Health AI, we're revolutionizing healthcare delivery through innovative AI-enabled solutions. As part of Amazon Health Services and One Medical, we're on a mission to make quality healthcare more accessible while improving patient outcomes. Our work directly impacts millions of lives by empowering patients and enabling healthcare providers to deliver more meaningful care. Role Overview: We're seeking an Applied Scientist to join our dynamic team in building state of the art AI/ML solutions for healthcare. This role offers a unique opportunity to work at the intersection of artificial intelligence and healthcare, developing solutions that will shape the future of medical services delivery. Key job responsibilities • Lead end-to-end development of AI/ML solutions for Amazon Health organization, including Amazon Pharmacy and One Medical • Research, design, and implement state-of-the-art machine learning models, with a focus on Large Language Models (LLMs) and Visual Language Models (VLMs) • Optimize and fine-tune models for production deployment, including model distillation for improved latency • Drive scientific innovation while maintaining a strong focus on practical business outcomes • Collaborate with cross-functional teams to translate complex technical solutions into tangible customer benefits • Contribute to the broader Amazon Health scientific community and help shape our technical roadmap
US, MA, Boston
The Artificial General Intelligence (AGI) team is seeking a dedicated, skilled, and innovative Applied Scientist with a robust background in machine learning, statistics, quality assurance, auditing methodologies, and automated evaluation systems to ensure the highest standards of data quality, to build industry-leading technology with Large Language Models (LLMs) and multimodal systems. Key job responsibilities As part of the AGI team, an Applied Scientist will collaborate closely with core scientist team developing Amazon Nova models. They will lead the development of comprehensive quality strategies and auditing frameworks that safeguard the integrity of data collection workflows. This includes designing auditing strategies with detailed SOPs, quality metrics, and sampling methodologies that help Nova improve performances on benchmarks. The Applied Scientist will perform expert-level manual audits, conduct meta-audits to evaluate auditor performance, and provide targeted coaching to uplift overall quality capabilities. A critical aspect of this role involves developing and maintaining LLM-as-a-Judge systems, including designing judge architectures, creating evaluation rubrics, and building machine learning models for automated quality assessment. The Applied Scientist will also set up the configuration of data collection workflows and communicate quality feedback to stakeholders. An Applied Scientist will also have a direct impact on enhancing customer experiences through high-quality training and evaluation data that powers state-of-the-art LLM products and services. A day in the life An Applied Scientist with the AGI team will support quality solution design, conduct root cause analysis on data quality issues, research new auditing methodologies, and find innovative ways of optimizing data quality while setting examples for the team on quality assurance best practices and standards. Besides theoretical analysis and quality framework development, an Applied Scientist will also work closely with talented engineers, domain experts, and vendor teams to put quality strategies and automated judging systems into practice.
US, CA, Santa Clara
Amazon Quick Suite is an enterprise AI platform that transforms how organizations work with their data and knowledge. Combining generative AI-powered search, deep research capabilities, intelligent agents and automations, and comprehensive business intelligence, Quick Suite serves tens of thousands of users. Our platform processes thousands of queries monthly, helping teams make faster, data-driven decisions while maintaining enterprise-grade security and governance. From natural language interactions with complex datasets to automated workflows and custom AI agents, Quick Suite is redefining workplace productivity at unprecedented scale. We are seeking a Data Scientist II to join our Quick Data team, focusing on evaluation and benchmarking data development for Quick Suite features, with particular emphasis on Research and other generative AI capabilities. Our mission is to engineer high-quality datasets that are essential to the success of Amazon Quick Suite. From human evaluations and Responsible AI safeguards to Retrieval-Augmented Generation and beyond, our work ensures that Generative AI is enterprise-ready, safe, and effective for users at scale. As part of our diverse team—including data scientists, engineers, language engineers, linguists, and program managers—you will collaborate closely with science, engineering, and product teams. We are driven by customer obsession and a commitment to excellence. Key job responsibilities In this role, you will leverage data-centric AI principles to assess the impact of data on model performance and the broader machine learning pipeline. You will apply Generative AI techniques to evaluate how well our data represents human language and conduct experiments to measure downstream interactions. Specific responsibilities include: * Design and develop comprehensive evaluation and benchmarking datasets for Quick Suite AI-powered features * Leverage LLMs for synthetic data corpora generation; data evaluation and quality assessment using LLM-as-a-judge settings * Create ground truth datasets with high-quality question-answer pairs across diverse domains and use cases * Lead human annotation initiatives and model evaluation audits to ensure data quality and relevance * Develop and refine annotation guidelines and quality frameworks for evaluation tasks * Conduct statistical analysis to measure model performance, identify failure patterns, and guide improvement strategies * Collaborate with ML scientists and engineers to translate evaluation insights into actionable product improvements * Build scalable data pipelines and tools to support continuous evaluation and benchmarking efforts * Contribute to Responsible AI initiatives by developing safety and fairness evaluation datasets About the team Why AWS? Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses. Inclusive Team Culture Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon conferences, inspire us to never stop embracing our uniqueness. Mentorship & Career Growth We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud. Hybrid Work We value innovation and recognize this sometimes requires uninterrupted time to focus on a build. We also value in-person collaboration and time spent face-to-face. Our team affords employees options to work in the office every day or in a flexible, hybrid work model near one of our U.S. Amazon offices.
US, CA, San Francisco
Amazon launched the AGI Lab to develop foundational capabilities for useful AI agents. We built Nova Act - a new AI model trained to perform actions within a web browser. The team builds AI/ML infrastructure that powers our production systems to run performantly at high scale. We’re also enabling practical AI to make our customers more productive, empowered, and fulfilled. In particular, our work combines large language models (LLMs) with reinforcement learning (RL) to solve reasoning, planning, and world modeling in both virtual and physical environments. Our lab is a small, talent-dense team with the resources and scale of Amazon. Each team in the lab has the autonomy to move fast and the long-term commitment to pursue high-risk, high-payoff research. We’re entering an exciting new era where agents can redefine what AI makes possible. We’d love for you to join our lab and build it from the ground up! Key job responsibilities This role will lead a team of SDEs building AI agents infrastructure from launch to scale. The role requires the ability to span across ML/AI system architecture and infrastructure. You will work closely with application developers and scientists to have a impact on the Agentic AI industry. We're looking for a Software Development Manager who is energized by building high performance systems, making an impact and thrives in fast-paced, collaborative environments. About the team Check out the Nova Act tools our team built on on nova.amazon.com/act
US, MA, Boston
**This is an experimental role to support a business pilot and can potentially span up to 12 months** Embark on a transformative journey as our Expert Consultant, where intellectual rigor meets technological innovation. As an Expert Consultant, you will blend your advanced analytical skills and domain expertise to provide strategic oversight to our human-in-the-loop and model-in-the-loop data pipelines. You will also provide mentorship and guidance to junior team members. Your responsibilities will ensure data excellence through strategic oversight of high-quality data output, while delivering expert consultation throughout the pipeline and fostering iterative development. This position directly impacts the effectiveness and reliability of our AI solutions by maintaining the highest standards of data quality throughout the development process while building capability within the broader team. Key job responsibilities • Serve as a trusted domain advisor to cross-functional teams, providing strategic direction and specialized problem-solving support • Champion domain knowledge sharing across multiple channels and teams to maintain data quality excellence and standardization • Drive collaborative efforts with science teams to optimize output of complex data collections in your domain expertise, ensuring data excellence through iterative feedback loops • Foster team excellence through mentorship and motivation of peers and junior team members • Make informed decisions on behalf of our customers, ensuring that selected code meets industry standards, best practices, and specific client needs • Collaborate with AI teams to innovate model-in-the-loop and human-in-the-loop approaches, to ensure the collection of high-quality data, safeguarding data privacy and security for LLM training, and more. • Stay abreast of the latest developments in how LLMs and GenAI can be applied to your area of expertise to ensure our evaluations remain cutting-edge. • Develop and write demonstrations to illustrate "what good data looks like" in terms of meeting benchmarks for quality and efficiency • Provide detailed feedback and explanations for your evaluations, helping to refine and improve the LLM's understanding and output
US, MA, Boston
**This is an experimental role to support a business pilot and can potentially span up to 12 months** Embark on a transformative journey as our Expert Consultant, where intellectual rigor meets technological innovation. As an Expert Consultant, you will blend your advanced analytical skills and domain expertise to provide strategic oversight to our human-in-the-loop and model-in-the-loop data pipelines. You will also provide mentorship and guidance to junior team members. Your responsibilities will ensure data excellence through strategic oversight of high-quality data output, while delivering expert consultation throughout the pipeline and fostering iterative development. This position directly impacts the effectiveness and reliability of our AI solutions by maintaining the highest standards of data quality throughout the development process while building capability within the broader team. Key job responsibilities • Serve as a trusted domain advisor to cross-functional teams, providing strategic direction and specialized problem-solving support • Champion domain knowledge sharing across multiple channels and teams to maintain data quality excellence and standardization • Drive collaborative efforts with science teams to optimize output of complex data collections in your domain expertise, ensuring data excellence through iterative feedback loops • Foster team excellence through mentorship and motivation of peers and junior team members • Make informed decisions on behalf of our customers, ensuring that selected code meets industry standards, best practices, and specific client needs • Collaborate with AI teams to innovate model-in-the-loop and human-in-the-loop approaches, to ensure the collection of high-quality data, safeguarding data privacy and security for LLM training, and more. • Stay abreast of the latest developments in how LLMs and GenAI can be applied to your area of expertise to ensure our evaluations remain cutting-edge. • Develop and write demonstrations to illustrate "what good data looks like" in terms of meeting benchmarks for quality and efficiency • Provide detailed feedback and explanations for your evaluations, helping to refine and improve the LLM's understanding and output
US, WA, Seattle
We are working on improving shopping on Amazon using the conversational capabilities of large language models and through customer behavioral data to make them more personalized for each customer. We are searching for pioneers who are passionate about technology, innovation, and customer experience, and are ready to make a lasting impact on the industry. In this role, you will be managing a team working on Large Language Model (LLM) and/or Vision-Language Model (VLM) post-training and alignment for new shopping experiences. You’ll be working with talented scientists, engineers, and technical program managers (TPM) to innovate on behalf of our customers. If you’re fired up about being part of a dynamic, driven team, then this is your moment to join us on this exciting journey!