Responsible AI in the wild: Lessons learned at AWS

Real-world deployment requires notions of fairness that are task relevant and responsive to the available data, recognition of unforeseen variation in the “last mile” of AI delivery, and collaboration with AI activists.

When we first joined AWS AI/ML as Amazon Scholars over three years ago, we had already been doing scientific research in the area now known as responsible AI for a while. We had authored a number of papers proposing mathematical definitions of fairness and machine learning (ML) training algorithms enforcing them, as well as methods for ensuring strong notions of privacy in trained models. We were well versed in adjacent subjects like explainability and robustness and were generally denizens of the emerging responsible-AI research community. We even wrote a general-audience book on these topics to try to explain their importance to a broader audience.

Related content
Generative AI raises new challenges in defining, measuring, and mitigating concerns about fairness, toxicity, and intellectual property, among other things. But work has started on the solutions.

So we were excited to come to AWS in 2020 to apply our expertise and methodologies to the ongoing responsible-AI efforts here — or at least, that was our mindset on arrival. But our journey has taken us somewhere quite different, somewhere more consequential and interesting than we expected. It’s not that the definitions and algorithms we knew from the research world aren’t relevant — they are — but rather that they are only one component of a complex AI workstream comprising data, models, services, enterprise customers, and end-users. It’s also a workstream in which AWS is uniquely situated due to its pioneering role in cloud computing generally and cloud AI services specifically.

Our time here has revealed to us some practical challenges of which we were previously unaware. These include diverse data modalities, “last mile” effects with customers and end-users, and the recent emergence of AI activism. Like many good interactions between industry and academia, what we’ve learned at AWS has altered our research agenda in healthy ways. In case it’s useful to anyone else trying to parse the burgeoning responsible-AI landscape (especially in the generative-AI era), we thought we’d detail some of our experiences here.

Modality matters

One of our first important practical lessons might be paraphrased as “modality matters”. By this we mean that the particular medium in which an AI service operates (such as visual images or spoken or written language) matters greatly in how we analyze and understand it from both performance and responsible-AI perspectives.

Consider specifically the desire for trained models be “fair”, or free of significant demographic bias. Much of the scientific literature on ML fairness assumes that the features used to compare performance across groups (which might include gender, race, age, and other attributes) are readily available, or can be accurately estimated, in both training and test datasets.

Related content
Two of the world’s leading experts on algorithmic bias look back at the events of the past year and reflect on what we’ve learned, what we’re still grappling with, and how far we have to go.

If this is indeed the case (as it might be for some spreadsheet-like “tabular” datasets recording things like medical or financial records, in which a person’s age and gender might be explicit columns), we can more easily test a trained model for bias. For instance, in a medical diagnosis application we might evaluate the model to make sure the error rates are approximately the same across genders. If these rates aren’t close enough, we can augment our data or retrain the model in various ways until the evaluation is passed to satisfaction.

But many cloud AI/ML services operate on data that simply does not contain explicit demographic information. Rather, these services live in entirely different modalities such as speech, natural language, and vision. Applications such as our speech recognition and transcription services take as input time series of frequencies that capture spoken utterances. Consequently, there are not direct annotations in the data of things like gender, race, or age.

But what can be more readily detected from speech data, and are also more directly related to performance, are regional dialects and accents — of which there are dozens in North American English alone. English-language speech can also feature non-native accents, influenced more by the first languages of the speakers than by the regions in which they currently live. This presents an even more diverse landscape, given the large number of first languages and the international mobility of speakers. And while spoken accents may be weakly correlated or associated with one or more ancestry groups, they are usually uninformative on things like age and gender (speakers with a Philadelphia accent may be young or old; male, female or nonbinary; etc.). Finally, the speech of even a particular person may exhibit many other sources of variation, such as situational stress and fatigue.

Regional dialects.jpeg
Data — such as regional variations in word choice and accents — may lead toward alternative notions of fairness that are more task-relevant, as with word error rates across dialects and accents.

What is the responsible-AI practitioner to do when confronted with so many different accents and other moving parts, in a task as complex as speech transcription? At AWS, our answer is to meet the task and data on their own terms, which in this case involves some heavy lifting: meticulously gathering samples from large populations of representative speakers with different accents and carefully transcribing each word. The “representative” is important here: while it might be more expedient to (for instance) gather this data from professional actors trained in diction, such data would not be typical of spoken language in the wild.

Related content
Both secure multiparty computation and differential privacy protect the privacy of data used in computation, but each has advantages in different contexts.

We also gather speech data that exhibits variability along other important dimensions, including the acoustic conditions during recording (varying amounts and types of background noise, recordings made via different mobile-phone handsets, whose microphones may vary in quality, etc.). The sheer number of combinations makes obtaining sufficient coverage challenging. (In some domains such as computer vision, coverage issues that are similar — variability across visual properties such as skin tone, lighting conditions, indoor vs. outdoor settings, and so on — have led to increased interest in synthetic data to augment human-generated data, including for fairness testing here at AWS.)

Once curated, such datasets can be used for training a transcription model that is not only good overall but also roughly equally performant across accents. And “performant” here means something more complex than in a simple prediction task; speech recognition typically uses a measure like the word error rate. On top of all the curation and annotations above, we also annotate some data by self-reported speaker demographics to make sure we’re fair not just by accent but by race and gender as well, as detailed in the service’s accompanying service card.

Our overarching point here is twofold. First, while as a society we tend to focus on dimensions such as race and gender when speaking about and assessing fairness, sometimes the data simply doesn’t permit such assessments, and it may not be a good idea to impute such dimensions to the data (for instance, by trying to infer race from speech signals). And second, in such cases the data may lead us toward alternative notions of fairness that might be more task-relevant, as with word error rates across dialects and accents.

The last mile of responsible AI

The specific properties of individuals that can or cannot (or should not) be gleaned from a particular dataset or modality are not the only things that may be out of the direct control of AI developers — especially in the era of cloud computing. As we have seen above, it’s challenging work to get coverage of everything you can anticipate. It’s even harder to anticipate everything.

The supply chain phrase “the last mile” refers to the fact that “upstream” providers of goods and products may have limited control over the “downstream” suppliers that directly connect to end-users or consumers. The emergence of cloud providers like AWS has created an AI service supply chain with its own last-mile challenges.

Related content
The team’s latest research on privacy-preserving machine learning, federated learning, and bias mitigation.

AWS AI/ML provides enterprise customers with API access to services like speech transcription because many want to integrate such services into their own workflows but don’t have the resources, expertise, or interest to build them from scratch. These enterprise customers sit between the general-purpose services of a cloud provider like AWS and the final end-users of the technology. For example, a health care system might want to provide cloud speech transcription services optimized for medical vocabulary to allow doctors to take verbal notes during their patient rounds.

As diligent as we are at AWS at battle-testing our services and underlying models for state-of-the-art performance, fairness, and other responsible-AI dimensions, it is obviously impossible to anticipate all possible downstream use cases and conditions. Continuing our health care example, perhaps there is a floor of a particular hospital that has new and specialized imaging equipment that emits background noise at a specific regularity and acoustic frequency. In the likely event that these exact conditions were not represented in either the training or test data, it’s possible that overall word error rates will not only be higher but may be so differentially across accents and dialects.

Such last-mile effects can be as diverse as the enterprise customers themselves. With time and awareness of such conditions, we can use targeted training data and customer-side testing to improve downstream performance. But due to the proliferation of new use cases, it is an ever-evolving process, not one that is ever “finished”.

AI activism: from bugs to bias

It’s not only cloud customers whose last miles may present conditions that differ from those during training and testing. We live in a (healthy) era of what might be called AI activism, in which not only enterprises but individual citizens — including scientists, journalists, and members of nonprofit organizations — can obtain API or open-source access to ML services and models and perform their own evaluations on their own curated datasets. Such tests are often done to highlight weaknesses of the technology, including shortfalls in overall performance and fairness but also potential security and privacy vulnerabilities. As such, they are typically performed without the AI developer’s knowledge and may be first publicized in both research and mainstream media outlets. Indeed, we have been on the receiving end of such critical publicity in the past.

Related content
Technique that mixes public and private training data can meet differential-privacy criteria while cutting error increase by 60%-70%.

To date, the dynamic between AI developers and activists has been somewhat adversarial: activists design and conduct a private experimental evaluation of a deployed AI model and report their findings in open forums, and developers are left to evaluate the claims and make any needed improvements to their technology. It is a dynamic that is somewhat reminiscent of the historical tensions between more traditional software and security developers and the ethical and unethical hacker communities, in which external parties probe software, operating systems, and other platforms for vulnerabilities and either expose them for the public good or exploit them privately for profit.

Over time the software community has developed mechanisms to alter these dynamics to be more productive than adversarial, in particular in the form of bug bounty programs. These are formal events or competitions in which software developers invite the hacker community to deliberately find vulnerabilities in their technology and offer financial or other rewards for reporting and describing them to the developers.

Bias bounties.png
In a fair-ML (“bias bounty”) competition, different teams (x-axis) focus on different demographic features (y-axis) in the dataset, indicating that crowdsourced bias mitigation can help contend with the breadth of possible sources of bias. (The darker the blue, the greater the use of the feature.)

In the last couple of years, the ideas and motivations behind bug bounties have been adopted and adapted by the AI development community, in the form of “bias bounties”. Rather than finding bugs in traditional software, participants are invited to help identify demographic or other biases in trained ML models and systems. Early versions of this idea were informal hackathons of short duration focused on finding subsets of a dataset on which a model underperformed. But more recent proposals incubated at AWS and elsewhere include variants that are more formal and algorithmic in nature. The explosion of models, interest in, and concerns about generative AI have also led to more codified and institutionalized responsible-AI methodologies such as the HELM framework for evaluating large language models.

We view these recent developments — AI developers opening up their technology and its evaluation to a wider community of stakeholders than just enterprise customers, and those stakeholders playing an active role in identifying necessary improvements in both technical and nontechnical ways — as healthy and organic, a natural outcome of the complex and evolving AI industry. Indeed, such collaborations are in keeping with our recent White House commitments to external testing and model red-teaming.

Responsible AI is neither a problem to be “solved” once and for all, nor a problem that can be isolated to a single location in the pipeline stretching from developers to their customers to end-users and society at large. Developers are certainly the first line where best practices must be established and implemented and responsible-AI principles defended. But the keys to the long-term success of the AI industry lie in community, communication, and cooperation among all those affected by it.

Related content

CN, 31, Shanghai
上海职位 - 如果希望在上海工作,请投递本职位。 毕业时间:2026年10月 - 2027年9月之间毕业的应届毕业生 · 投递须知: 1 填写简历申请时,请把必填和非必填项都填写完整。提交简历之后就无法修改了哦! 2 学校的英文全称请准确填写。中英文对应表,请点击链接查看 https://docs.qq.com/sheet/DVmdaa1BCV0RBbnlR?tab=BB08J2 3 简历不限中英文。 如果您正在攻读自然语言处理(NLP)、信息检索(IR)、机器学习、生成式人工智能或相关方向的硕士或博士学位,并希望将前沿科学研究转化为服务真实客户的产品,我们诚挚邀请您加入亚马逊 International Technology 搜索团队。 我们的目标是帮助亚马逊客户更准确地找到所需商品,并发现符合其需求和兴趣的新商品。您每天的工作都将直接影响全球数百万客户的购物体验。团队使用 TB 级商品、查询和客户行为数据,持续推进搜索、推荐、自然语言理解以及生成式 AI 技术的发展。 在这个岗位中,您将研究并应用 NLP、IR、深度学习、大语言模型(LLM)和基础模型等前沿技术,解决搜索理解、相关性排序、语义匹配、个性化和对话式购物等问题。您将有机会探索预训练、监督微调(SFT)、参数高效微调、检索增强生成(RAG)、提示优化和智能体(Agent)等技术,并针对业务场景建立可靠的离线与在线评估方法。 您将与应用科学家、软件工程师和产品经理密切合作,完成从问题定义、数据分析、算法设计和实验验证,到模型部署、在线测试和持续迭代的完整闭环。您需要根据客户价值和业务目标选择合适的技术方案,并在模型质量、可靠性、安全性、推理延迟和计算成本之间做出合理权衡。 Key job responsibilities Key job responsibilities · 针对 Amazon 搜索和购物体验中的实际问题,提出可验证的科学假设,设计并实现机器学习、NLP、IR 或 LLM 解决方案。 · 使用大规模商品、查询和客户行为数据训练、微调和评估模型,建立可重复的实验与评估流程。 · 探索基础模型在搜索、推荐和对话式购物中的应用,包括 RAG、模型微调、提示优化和 Agent 等方向。 · 设计覆盖相关性、事实性、鲁棒性、安全性、延迟和成本的评估指标,并通过离线实验、A/B 测试和客户反馈验证效果。 · 与工程和产品团队合作,将原型转化为可扩展、可维护的生产系统,并持续分析和改进线上表现。 · 跟踪学术界和工业界的最新进展,形成技术文档,并在适当情况下向内部或外部科学社区分享研究成果。 基本要求 · 正在攻读或已获得计算机科学、计算机工程、机器学习、人工智能、运筹学、统计学或相关领域的硕士或博士学位。 · 具备机器学习或深度学习的基础知识,以及实验设计、统计分析和模型评估经验。 · 具备使用代码和工具实现、训练和评估算法的经验。 · 至少熟练使用一种编程语言,例如 Python、Java 或 C++。 · 了解 NLP、IR、推荐系统或生成式 AI 中至少一个方向的基本方法。 优先条件 · 在 NLP、IR、机器学习、数据挖掘或生成式 AI 相关顶级会议或期刊发表过论文,或有高质量研究项目经历。 · 熟悉 Transformer、LLM 或基础模型,并具有预训练、监督微调(SFT)、参数高效微调、偏好优化或推理优化中的一种或多种实践经验。 · 具有 RAG、向量检索、Embedding、语义匹配、Agent 或工具调用系统的研究或开发经验。 · 熟悉 PyTorch、TensorFlow 等深度学习框架,以及 Hugging Face Transformers 等常用 LLM 工具链。 · 具有搜索引擎或推荐系统经验,尤其是在索引、召回、排序、查询理解、个性化或在线实验方面。 · 具有 LLM 评估经验,能够从相关性、事实性、幻觉、鲁棒性、安全性、延迟和成本等维度衡量系统质量。 · 具有大规模数据处理、分布式训练、模型压缩或高效推理经验。 · 具备良好的批判性思维和技术沟通能力,能够清楚地解释模型选择、实验结果及其局限性,并与跨职能团队合作解决开放性问题。
IN, KA, Bengaluru
If you have ever bought or sold anything on Amazon, you have touched Amazon Marketplace. Amazon’s Marketplace business is one of the largest in the world. We are now in 23 countries. We are growing fast, with customers in many more countries. Amazon’s platform is the engine that powers Amazon’s Marketplace businesses, and Sellers rely on this platform and our support to start selling on Amazon and to grow their business. Amazon Marketplace enables millions of Sellers worldwide to list hundreds of millions of products and manage orders for inventory across dozens of different categories and languages. While working with millions of Sellers worldwide, we constantly strive to improve the selection for Customers and the capabilities of our platform for Sellers. The Seller Fulfillment Services (SFS) team is looking for a motivated and innovative Data Scientist with strong analytical skills and practical experience to join our science team. As a key member of the SFS science team, you will provide expertise that helps accelerate the business. You will build science solutions that will help us to provide our customers with the largest selection of merchants at the lowest, and the most reliable delivery service regardless of the seller. You will research, design and improve on the models that will impact Amazon’s customer directly. You will be working in a highly collaborative environment partnering with various science, product management, engineering, operations, finance, business intelligence and analytics teams to develop science models to solve business problems. You will need to understand the business requirements and translate them into complex analytical outputs. You will design tests to explain performance of the models from impact on customer and cost perspective. You will create ML models to capture features impacting performance. You should be comfortable building prototypes, testing and improving them given the feedback from the real time data. You should be able to present your model and findings to a various range of stakeholders. Looking for candidate with expertise in the areas of machine learning. The candidate will be expected to work on numerous aspects, such as feature engineering, modeling, and hyper-parameter tuning. Challenges will involve dealing with very large data sets and requirements on throughput. Key job responsibilities Design, implement, test, deploy, and maintain innovative science solutions to accelerate our business. Create experiments and prototype implementations of new learning algorithms and prediction techniques. Collaborate with scientists, engineers, product managers, and stakeholders to design and implement software solutions for science problems. Use best practices to ensure a high standard of quality for all of the team deliverables
IN, KA, Bengaluru
RBS (Retail Business Services) Tech team works towards enhancing the customer experience (CX) and their trust in product data by providing technologies to find and fix Amazon CX defects at scale. Our platforms help in improving the CX in all phases of customer journey, including selection, discoverability & fulfilment, buying experience and post-buying experience (product quality and customer returns). The team also develops GenAI platforms for automation of Amazon Stores Operations. As a Sciences team in RBS Tech, we focus on foundational ML research and develop scalable state-of-the-art ML solutions to solve the problems covering customer experience (CX) and Selling partner experience (SPX). We work to solve problems related to multi-modal understanding (text and images), task automation through multi-modal LLM Agents, supervised and unsupervised techniques, multi-task learning, multi-label classification, aspect and topic extraction for Customer Anecdote Mining, image and text similarity and retrieval using NLP and Computer Vision for product groupings and identifying duplicate listings in product search results. Key job responsibilities As an Applied Scientist, you will be responsible to design and deploy scalable GenAI, NLP and Computer Vision solutions that will impact the content visible to millions of customer and solve key customer experience issues. You will develop novel LLM, deep learning and statistical techniques for task automation, text processing, image processing, pattern recognition, and anomaly detection problems. You will define the research and experiments strategy with an iterative execution approach to develop AI/ML models and progressively improve the results over time. You will partner with business and engineering teams to identify and solve large and significantly complex problems that require scientific innovation. You will independently file for patents and/or publish research work where opportunities arise. The RBS org deals with problems that are directly related to the selling partners and end customers and the ML team drives resolution to organization level problems. Therefore, the Applied Scientist role will impact the large product strategy, identifies new business opportunities and provides strategic direction which is very exciting.
US, NY, New York
We are looking for an Applied Scientist III to set the scientific direction for the next generation of agentic AI applications that guide Amazon advertisers. In this role you will define, lead and build the science behind agentic systems that reason, plan, and act autonomously to manage and optimize ad campaigns based on a deep understanding of the advertiser and the marketplace. You will own the agentic architecture end to end, partnering closely with product and engineering leaders to translate a long-term science vision into concrete research and engineering roadmaps. Working backwards from the needs of millions of advertisers, you will take the lead on medium-to-large, ambiguous problems where neither the problem nor the solution is well defined, and deliver customer-facing products that help advertisers create, optimize, and grow their campaigns. You will invent new methods at the product level, and drive their adoption across multiple teams. This role combines science leadership, technical depth, product focus, and business understanding: you will raise the science bar, build consensus on approach across partners, and mentor scientists and engineers while remaining deeply hands-on with the hardest technical problems. Key job responsibilities As an Applied Scientist III on this team you will: - Define the science vision for the agentic campaign management system and, with product and engineering leaders, turn it into delivery roadmaps. - Build agentic systems that autonomously manage and optimize ad campaigns — encoding auction and marketplace dynamics (bidding, budget pacing, keyword and targeting decisions) while balancing advertiser ROI, shopper experience, and marketplace health. - Define and curate the datasets and signals needed to train and evaluate these agents — advertiser and campaign data, auction and bid/budget signals, impressions, clicks, conversions, and search-term/keyword performance. - Stay deeply hands-on: write production-quality, critical-path code and build core components that take agentic systems from prototype to launch. - Own the agentic architecture — planning, tool use and integration (e.g., MCP), long-horizon reasoning (e.g., ReAct, CoT/ToT), and multi-agent orchestration — and stay deeply hands-on, writing production-quality, critical-path code from prototype to launch. - Define the evaluation and safety methodology for agent workflows and drive its adoption as the bar for reliability and trust. - Drive the team's scientific agenda, mentor scientists and engineers, and represent the team in the internal and external scientific community. About the team The Sponsored Products and Brands team at Amazon Ads is re-imagining the advertising landscape through the latest generative AI technologies, revolutionizing how millions of customers discover products and engage with brands across Amazon.com and beyond. We are at the forefront of re-inventing advertising experiences, bridging human creativity with artificial intelligence to transform every aspect of the advertising lifecycle from ad creation and optimization to performance analysis and customer insights. We are a passionate group of innovators dedicated to developing responsible and intelligent AI technologies that balance the needs of advertisers, enhance the shopping experience, and strengthen the marketplace. If you're energized by solving complex challenges and pushing the boundaries of what's possible with AI, join us in shaping the future of advertising. This team within Sponsored Products and Brands is focused on guiding and supporting millions of advertisers to meet their advertising needs of creating and managing ad campaigns. At this scale, the complexity of diverse advertiser goals, campaign types, and market dynamics creates both a massive technical challenge and a transformative opportunity: even small improvements in guidance systems can have outsized impact on advertiser success and Amazon’s retail ecosystem. Our vision is to build a highly personalized, context-aware agentic advertiser guidance system that leverages LLMs together with tools such as auction simulations, ML models, and optimization algorithms. This agentic framework, will operate across both chat and non-chat experiences in the ad console, scaling to natural language queries as well as autonomously manage campaigns based on deep understanding of the advertiser. To execute this vision, we collaborate closely with stakeholders across Ad Console, Sales, and Marketing to identify opportunities—from high-level product guidance down to granular keyword recommendations—and deliver them through a tailored, personalized experience. Our work is grounded in state-of-the-art agent architectures, tool integration, reasoning frameworks, and model customization approaches (including tuning, MCP, and preference optimization), ensuring our systems are both scalable and adaptive.
US, WA, Seattle
Innovators wanted! Are you an entrepreneur? A builder? A dreamer? This role is part of an Amazon Special Projects team that takes the company's Think Big leadership principle to the next level. We focus on creating entirely new products and services with a goal of positively impacting the lives of our customers. No industries or subject areas are out of bounds. If you're interested in innovating at scale to address big challenges in the world, this is the team for you. As a Research Scientist, you will work with a unique and gifted team developing exciting products for consumers and collaborate with cross-functional teams. Our team rewards intellectual curiosity while maintaining a laser-focus in bringing products to market. Competitive candidates are responsive, flexible, and able to succeed within an open, collaborative, entrepreneurial, startup-like environment. At the intersection of both academic and applied research in this product area, you have the opportunity to work together with some of the most talented scientists, engineers, and product managers. Here at Amazon, we embrace our differences. We are committed to furthering our culture of inclusion. We have thirteen employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We are constantly learning through programs that are local, regional, and global. Amazon's culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Our team highly values work-life balance, mentorship and career growth. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We care about your career growth and strive to assign projects and offer training that will challenge you to become your best.
US, WA, Seattle
Prime Video is a first-stop entertainment destination offering customers a vast collection of premium programming in one app available across thousands of devices. Prime members can customize their viewing experience and find their favorite movies, series, documentaries, and live sports – including Amazon MGM Studios-produced series and movies; licensed fan favorites; and programming from Prime Video add-on subscriptions such as Apple TV+, Max, Crunchyroll and MGM+. All customers, regardless of whether they have a Prime membership or not, can rent or buy titles via the Prime Video Store, and can enjoy even more content for free with ads. Prime Video Personalization and Discovery organization operates at a global scale and is a critical part of Prime Video’s flywheel. We are responsible for matching customers with the right content at the right time, at all touch points throughout customers' content discovery journey: from home page, title page, and search page, to showing recommendations at the end of an episode or a movie. We drive both immediate and long-term business objectives while delighting our customers. Key job responsibilities As a Sr. Manager, Applied Science, in the Prime Video Personalization and Discovery organization, you will be responsible for optimizing the complete customer experience, across the touch points throughout customers’ discovery journey. This includes building AI and optimization solutions, working with product, engineering teams to deliver the optimal balance of customer delight and business outcomes. About the team Prime Video Personalization and Discovery (PVPD) is dedicated to creating a highly personalized content discovery experience that not only delights our customers but also drives both short-and long-term business goals. Our scope includes personalized recommendations, search, marketing, and the advanced machine learning technology and infrastructure that underpins these experiences. Our mission is to automate and enhance customer engagement through personalization, using ML and Generative AI.
US, NY, New York
Application deadline: Applications will be accepted on an ongoing basis Applied Scientists in AWS Automated Reasoning develop and apply bleeding-edge formal methods, automated reasoning techniques, and neurosymbolic approaches to ensure the security, reliability, and correctness of Amazon/AWS services and customer applications. Our tools are called billions of times daily, powering the backbone of Amazon's products and services. We are changing the way computer systems are developed and operated, raising the bar for security, durability, availability, and quality. At Amazon, automated reasoning is central to maintaining customer trust and delivering delightful customer experiences. Application areas span cloud infrastructure verification, cryptographic assurance, AI safety, drone safety, and formal guarantees for generative AI systems. Our methods range from interactive theorem proving and constraint solving to neuro-inspired proof search This is a unique opportunity to get in early on a fast-growing segment of the business and help shape the technology, product, and business. You will have a chance to utilize your deep technical expertise within a fast-moving environment and make a large business and customer impact. Key job responsibilities • Design and implement algorithms and formal methods for automated reasoning, including constraint solving, model checking, static analysis, theorem proving, and program synthesis to verify the correctness, security, and reliability of computing systems. • Solve large or significantly complex problems that require deep knowledge and scientific innovation in your domain; own strategic problem solving and take the lead on design, implementation, and delivery of solutions with long-term quantifiable impact. • Develop new decision procedures, heuristics, and search strategies that improve the scalability and accuracy of verification tools; build and deploy production-grade automated reasoning systems at Amazon scale. • Explore and apply generative AI and machine learning techniques to enhance automated reasoning capabilities, including learning-based heuristics for search and optimization, neural approaches to symbolic reasoning, and methods for verifying the correctness of AI-generated code. • Develop automated reasoning techniques for generative AI and agentic coding systems, including methods for ensuring the safety and alignment of autonomous software agents and applying formal guarantees to large language model outputs. • Conduct original research snd publish findings in peer-reviewed venues. • Work with customer teams to understand the nature of their software and the properties they need to establish; identify tools and methods capable of addressing verification needs, including novel analysis capabilities. • Provide cross-organizational technical influence, increasing productivity and effectiveness by sharing deep knowledge and experience; collaborate with partner teams to translate verification capabilities into production systems. • Mentor scientists and engineers on formal methods, neurosymbolic techniques, and best practices for building reliable automated reasoning systems; assist in career development of others.
US, CA, Santa Clara
The Data Intelligence team is a new function within Amazon Customer Service (CS). We own the end-to-end process of defining, building, implementing, and monitoring a comprehensive data strategy. We also develop and apply Generative Artificial Intelligence (GenAI), Machine Learning (ML), Ontology, and Natural Language Processing (NLP) to enhance customer service associate and customer experiences. As an Applied Scientist, you'll own the definition and implementation of customer-focused, AI-driven innovation in Amazon Customer Service globally, leveraging GenAI, ML, and/or NLP to transform complex business requirements and customer needs into innovative technology solutions. Your expertise will be key in shaping data-driven strategies and addressing complex data challenges. With your expertise in AI, text analysis, embeddings, language modeling, and generation, you'll design and develop scalable AI-powered technology solutions, prioritize initiatives, drive data-driven insights, and deliver business impact. This position will advance applied science best practices, leverage data and AI to drive customer experience improvements, and set new global standards for customer experience. This role requires you to work with a cross-functional team, including scientists, engineers, and product managers, to develop scalable and maintainable AI solutions for both structured and unstructured data. The ideal candidate has strong technical skills in AI techniques (e.g., automated reasoning, reasoning, planning, knowledge representation), excellent written documentation skills, and experience with big data technologies. Success in this role requires combining deep business knowledge with hands-on technical skills to solve customer problems and address complex technical challenges. Key job responsibilities - Develop innovative solutions to complex problems (e.g., Automated Reasoning for Trusted AI-Enabled Customer Service). - Apply technical expertise to implement novel algorithms and modeling solutions, in collaboration with other scientists and engineers. - Analyze data and define metrics to identify actionable insights and measure improvements in customer experience. - Communicate results and insights to both technical and non-technical audiences through written reports, presentations, and internal/external publications. - Collaborate with product management and engineering teams to integrate and optimize models in production systems. A day in the life A typical day as an Applied Scientist in the Data Intelligence team involves combining business expertise with hands-on problem-solving in ML and AI. The role encompasses tackling complex data initiatives, ensuring alignment with customer needs and business objectives, and translating business requirements into practical AI-driven solutions. Working collaboratively with cross-functional teams, this position involves designing and enhancing AI models, focusing on efficiency, precision, and scalability. Daily activities include ensuring data quality, monitoring model performance, and generating actionable insights from vast amounts of information. Each day presents opportunities to resolve complex technical challenges, advance important AI projects, and conceive innovative ways to leverage data in transforming the customer experience. About the team The Data Intelligence team is a new function within Amazon Customer Service. We develop and apply Generative Artificial Intelligence (GenAI), Machine Learning (ML), and Natural Language Processing (NLP) techniques to enhance customer service associate and customer experiences.
US, CA, East Palo Alto
As part of the AWS Solutions organization, we have a vision to provide business applications, leveraging Amazon’s unique experience and expertise, that are used by millions of companies worldwide to manage day-to-day operations. We will accomplish this by accelerating our customers’ businesses through delivery of intuitive and differentiated technology solutions that solve enduring business challenges. We blend vision with curiosity and Amazon’s real-world experience to build opinionated, turnkey solutions. Where customers prefer to buy over build, we become their trusted partner with solutions that are no-brainers to buy and easy to use. Key job responsibilities Everyone on the team needs to be entrepreneurial, wear many hats and work in a highly collaborative environment that’s more startup than big company. We’ll need to tackle problems that span a variety of domains: computer vision, image recognition, machine learning, real-time and distributed systems. As an Applied Scientist, you will help solve a variety of technical challenges and mentor other scientists. You will tackle challenging, novel situations every day and given the size of this initiative, you’ll have the opportunity to work with multiple technical teams at Amazon in different locations. You should be comfortable with a degree of ambiguity that’s higher than most projects and relish the idea of solving problems that, frankly, haven’t been solved at scale before - anywhere. Along the way, we guarantee that you’ll learn a ton, have fun and make a positive impact on millions of people. A key focus of this role will be developing and implementing advanced visual reasoning systems that can understand complex spatial relationships and object interactions in real-time. You'll work on designing autonomous AI agents that can make intelligent decisions based on visual inputs, understand customer behavior patterns, and adapt to dynamic retail environments. This includes developing systems that can perform complex scene understanding, reason about object permanence, and predict customer intentions through visual cues. About the team Just Walk Out (JWO) is a new kind of store with no lines and no checkout—you just grab and go! Customers simply use the Amazon Go app to enter the store, take what they want from our selection of fresh, delicious meals and grocery essentials, and go! Our checkout-free shopping experience is made possible by our Just Walk Out Technology, which automatically detects when products are taken from or returned to the shelves and keeps track of them in a virtual cart. When you’re done shopping, you can just leave the store. Shortly after, we’ll charge your account and send you a receipt. Check it out at amazon.com/go. Designed and custom-built by Amazonians, our Just Walk Out Technology uses a variety of technologies including computer vision, sensor fusion, and advanced machine learning. Innovation is part of our DNA! Our goal is to be Earths’ most customer centric company and we are just getting started. We need people who want to join an ambitious program that continues to push the state of the art in computer vision, machine learning, distributed systems and hardware design.
US, CA, San Francisco
Join the next revolution in robotics at Amazon's Frontier AI & Robotics team, where you'll work alongside world-renowned AI pioneers to push the boundaries of what's possible in robotic intelligence. As a Member of Technical Staff, you'll be at the forefront of developing breakthrough foundation models that enable robots to perceive, understand, and interact with the world in unprecedented ways. You'll drive independent research initiatives in areas such as perception, manipulation, science understanding, locomotion, manipulation, sim2real transfer, multi-modal foundation models and multi-task robot learning, designing novel frameworks that bridge the gap between state-of-the-art research and real-world deployment at Amazon scale. In this role, you'll balance innovative technical exploration with practical implementation, collaborating with platform teams to ensure your models and algorithms perform robustly in dynamic real-world environments. You'll have access to Amazon's vast computational resources, enabling you to tackle ambitious problems in areas like very large multi-modal robotic foundation models and efficient, promptable model architectures that can scale across diverse robotic applications. Key job responsibilities - Drive independent research initiatives across the robotics stack, including robotics foundation models, focusing on breakthrough approaches in perception, and manipulation, for example open-vocabulary panoptic scene understanding, scaling up multi-modal LLMs, sim2real/real2sim techniques, end-to-end vision-language-action models, efficient model inference, video tokenization - Design and implement novel deep learning architectures that push the boundaries of what robots can understand and accomplish - Lead full-stack robotics projects from conceptualization through deployment, taking a system-level approach that integrates hardware considerations with algorithmic development, ensuring robust performance in production environments - Collaborate with platform and hardware teams to ensure seamless integration across the entire robotics stack, optimizing and scaling models for real-world applications - Contribute to the team's technical strategy and help shape our approach to next-generation robotics challenges A day in the life - Design and implement novel foundation model architectures and innovative systems and algorithms, leveraging our extensive infrastructure to prototype and evaluate at scale - Collaborate with our world-class research team to solve complex technical challenges - Lead technical initiatives from conception to deployment, working closely with robotics engineers to integrate your solutions into production systems - Participate in technical discussions and brainstorming sessions with team leaders and fellow scientists - Leverage our massive compute cluster and extensive robotics infrastructure to rapidly prototype and validate new ideas - Transform theoretical insights into practical solutions that can handle the complexities of real-world robotics applications About the team At Frontier AI & Robotics, we're not just advancing robotics – we're reimagining it from the ground up. Our team is building the future of intelligent robotics through innovative foundation models and end-to-end learned systems. We tackle some of the most challenging problems in AI and robotics, from developing sophisticated perception systems to creating adaptive manipulation strategies that work in complex, real-world scenarios. What sets us apart is our unique combination of ambitious research vision and practical impact. We leverage Amazon's massive computational infrastructure and rich real-world datasets to train and deploy state-of-the-art foundation models. Our work spans the full spectrum of robotics intelligence – from multimodal perception using images, videos, and sensor data, to sophisticated manipulation strategies that can handle diverse real-world scenarios. We're building systems that don't just work in the lab, but scale to meet the demands of Amazon's global operations. Join us if you're excited about pushing the boundaries of what's possible in robotics, working with world-class researchers, and seeing your innovations deployed at unprecedented scale.