Lessons learned from 10 years of DynamoDB

Prioritizing predictability over efficiency, adapting data partitioning to traffic, and continuous verification are a few of the principles that help ensure stability, availability, and efficiency.

Amazon DynamoDB is one of the most popular NoSQL database offerings on the Internet, designed for simplicity, predictability, scalability, and reliability. To celebrate DynamoDB’s 10th anniversary, the DynamoDB team wrote a paper describing lessons we’d learned in the course of expanding a fully managed cloud-based database system to hundreds of thousands of customers. The paper was presented at this year’s USENIX ATC conference.

The paper captures the following lessons that we have learned over the years:

  • Designing systems for predictability over absolute efficiency improves system stability. While components such as caches can improve performance, they should not introduce bimodality, in which the system has two radically different ways of responding to similar requests (e.g., one for cache misses and one for cache hits). Consistent behaviors ensure that the system is always provisioned to handle the unexpected. 
  • Adapting to customers’ traffic patterns to redistribute data improves customer experience. 
  • Continuously verifying idle data is a reliable way to protect against both hardware failures and software bugs in order to meet high durability goals. 
  • Maintaining high availability as a system evolves requires careful operational discipline and tooling. Mechanisms such as formal proofs of complex algorithms, game days (chaos and load tests), upgrade/downgrade tests, and deployment safety provide the freedom to adjust and experiment with the code without the fear of compromising correctness. 
Related content
Amazon DynamoDB was introduced 10 years ago today; one of its key contributors reflects on its origins, and discusses the 'never-ending journey' to make DynamoDB more secure, more available and more performant.

Before we dig deeper into these topics, a little terminology. A DynamoDB table is a collection of items (e.g., products), and each item is a collection of attributes (e.g., name, price, category, etc.). Each item is uniquely identified by its primary key. In DynamoDB, tables are typically partitioned, or divided into smaller sub-tables, which are assigned to nodes. A node is a set of dedicated computational resources — a virtual machine — running on a single server in a datacenter.

DynamoDB stores three copies of each partition, in different availability zones. This makes the partition highly available and durable because the availability zones’ storage resources share nothing and are substantially independent. For instance, we wouldn’t assign a partition and one of its copies to nodes that share a power supply, because a power outage would take both of them offline. The three copies of the same partition are known as a replication group, and there is a leader for the group that is responsible for replicating all the customer mutations and serving strongly consistent reads.

DynamoDB architecture.png
The DynamoDB architecture, including a request router, the partition metadata system, and storage nodes in different availability zones (AZs).

Those definitions in hand, let’s turn to our lessons learned.

Predictability over absolute efficiency

DynamoDB employs a lot of metadata caches in order to reduce latency. One of those caches stores the routing metadata for data requests. This cache is deployed on a fleet of thousands of request routers, DynamoDB’s front-end service.

In the original implementation, when the request router received the first request for a table, it downloaded the routing information for the entire table and cached it locally. Since the configuration information about partition replicas rarely changed, the cache hit rate was approximately 99.75%.

Related content
How Alexa scales machine learning models to millions of customers.

This was an amazing hit rate. However, on the flip side, the fallback mechanism for this cache was to hit the metadata table directly. When the cache becomes ineffective, the metadata table needs to instantaneously scale from handling 0.25% of requests to 100%. The sudden increase in traffic can cause the metadata table to fail, causing cascading failure in other parts of the system. To mitigate against such failures, we redesigned our caches to behave predictably.

First, we built an in-memory datastore called MemDS, which significantly reduced request routers’ and other metadata clients’ reliance on local caches. MemDS stores all the routing metadata in a highly compressed manner and replicates it across a fleet of servers. MemDS scales horizontally to handle all incoming requests to DynamoDB.

Second, we deployed a new local cache that avoids the bimodality of the original cache. All requests, even if satisfied by the local cache, are asynchronously sent to the MemDS. This ensures that the MemDS fleet is always serving a constant volume of traffic, regardless of cache hit or miss. The regular exercise of the fallback code helps prevent surprises during fallback.

DDB-MemDS.png
DynamoDB architecture with MemDS.

Unlike conventional local caches, MemDS sees traffic that is proportional to the customer traffic seen by the service; thus, during cache failures, it does not see a sudden amplification of traffic. Doing constant work removed the need for complex logic to handle edge cases around cache misses and reduced the reliance on local caches, improving system stability.

Reshaping partitioning based on traffic

Partitions offer a way to dynamically scale both the capacity and performance of tables. In the original DynamoDB release, customers explicitly specified the throughput that a table required in terms of read capacity units (RCUs) and write capacity units (WCUs). The original system assigned partitions to nodes based on both available space and computational capacity.

Related content
Optimizing placement of configuration data ensures that it’s available and consistent during “network partitions”.

As the demands on a table changed (because it grew in size or because the load increased), partitions could be further split to allow the table to scale elastically. Partition abstraction proved really valuable and continues to be central to the design of DynamoDB.

However, the early version of DynamoDB assigned both space and capacity to individual partitions on the basis of size, evenly distributing computational resources across table entries. This led to challenges of “hot partitions” and throughput dilution.

Hot partitions happened because customer workloads were not uniformly distributed and kept hitting a subset of items. Throughput dilution happened when partitions that had been split to handle increased load ended up with so few keys that they could quickly max out their meager allocated capacity.

Our initial response to these challenges was to add bursting and adaptive capacity (along with other features such as split for consumption) to DynamoDB. This line of work also led to the launch of on-demand tables.

Bursting is a way to absorb temporal spikes in workloads at a partition level. It’s based on the observation that not all partitions hosted by a storage node use their allocated throughput simultaneously.

Related content
Amazon researchers describe new method for distributing database tables across servers.

The idea is to let applications tap into unused capacity at a partition level on a best-effort basis to absorb short-lived spikes. DynamoDB still maintains workload isolation by ensuring that a partition can burst only if there is unused throughput at the node level.

DynamoDB also launched adaptive capacity to handle long-lived spikes that cannot be absorbed by the burst capacity. Adaptive capacity monitors traffic patterns and repartitions tables so that heavily accessed items reside on different nodes.

Both bursting and adaptive capacity had limitations, however. Bursting was helpful only for short-lived spikes in traffic, and it was dependent on nodes’ having enough throughput to support it. Adaptive capacity was reactive and kicked in only after transmission rates had been throttled down to avoid overloads.

To address these limitations, the DynamoDB team replaced adaptive capacity with global admission control (GAC). GAC builds on the idea of token buckets, in which bandwidth is allocated to network nodes as tokens, and the nodes “cash in” tokens in order to transmit data. Each request router maintains a local token bucket and communicates with GAC to replenish tokens at regular intervals (on the order of every few seconds). For an extra layer of defense, DynamoDB also uses token buckets at the partition level.

Continuous verification 

To provide durability and crash recovery, DynamoDB uses write-ahead logs, which record data writes before they occur. In the event of a crash, DynamoDB can use the write-ahead logs to reconstruct lost data writes, bringing partitions up to date.

Write-ahead logs are stored in all three replicas of a partition. For higher durability, the write-ahead logs are periodically archived to S3, an object store that is designed for more than 99.99% (in fact, 11 nines) durability. Each replica contains the most recent write-ahead logs, which are usually waiting to be archived. The unarchived logs are typically a few hundred megabytes in size.

Storage replica vs. log replica.png
Healing a storage replica by copying the B-tree can take several minutes, while adding a log replica, which takes only a few seconds, ensures that there is no impact on durability.

DynamoDB continuously verifies data at rest. Our goal is to detect any silent data errors or “bit rot” — bit errors caused by degradation of the storage medium. An example of continuous verification is the scrub process.

The scrub process verifies two things: that all three copies in a replication group have the same data and that the live replicas match a reference replica built offline using the archived write-ahead-log entries.

The verification is done by computing the checksum of the live replica and matching that with a snapshot of the reference replica. A similar technique is used to verify replicas of global tables. Over the years, we have learned that continuous verification of data at rest is the most reliable method of protecting against hardware failures, silent data corruption, and even software bugs.

Availability

DynamoDB regularly tests its resilience to node, rack, and availability zone (AZ) failures. For example, to test the availability and durability of the overall service, DynamoDB performs power-off tests. Using realistic simulated traffic, a job scheduler powers off random nodes. At the end of all the power-off tests, the test tools verify that the data stored in the database is logically valid and not corrupted.

Related content
Amazon Athena reduces query execution time by 14% by eliminating redundant operations.

The first point about availability is that it needs to be measurable. DynamoDB is designed for 99.999% availability for global tables and 99.99% availability for regional tables. To ensure that these goals are being met, DynamoDB continuously monitors availability at the service and table levels. The tracked availability data is used to estimate customer-perceived availability trends and trigger alarms if the number of errors that customers see crosses a certain threshold.

These alarms are called customer-facing alarms (CFAs). The goal of these alarms is to report any availability-related problems and proactively mitigate them either automatically or through operator intervention. The key point to note here is that availability is measured not only on the server side but on the client side.

We also use two sets of clients to measure the user-perceived availability. The first set of clients is internal Amazon services using DynamoDB as the data store. These services share the availability metrics for DynamoDB API calls as observed by their software.

The second set of clients is our DynamoDB canary applications. These applications are run from every AZ in the region, and they talk to DynamoDB through every public endpoint. Real application traffic allows us to reason about DynamoDB availability and latencies as seen by our customers. The canary applications offer a good representation of what our customers might be experiencing both long and short term.

The second point is that read and write availability need to be handled differently. A partition’s write availability depends on the health of its leader and of its write quorum, meaning two out of the three replicas from different AZs. A partition remains available as long as there are enough healthy replicas for a write quorum and a leader.

Related content
“Anytime query” approach adapts to the available resources.

In a large service, hardware failures such as memory and disk failures are common. When a node fails, all replication groups hosted on the node are down to two copies. The process of healing a storage replica can take several minutes because the repair process involves copying the B-tree — a data structure that maps partitions to storage locations — and write-ahead logs.

Upon detecting an unhealthy storage replica, the leader of a replication group adds a log replica to ensure there is no impact on durability. Adding a log replica takes only a few seconds, because the system has to copy only the most recent write-ahead logs from a healthy replica; reconstructing the more memory-intensive B-tree can wait. Quick healing of affected replication groups using log replicas thus ensures the high durability of the most recent writes. Adding a log replica is the fastest way to ensure that the write quorum of the group is always met. This minimizes disruption to write availability due to an unhealthy write quorum. The leader replica serves consistent reads.

Introducing log replicas was a big change to the system, but the Paxos consensus protocol, which is formally provable, gave us the confidence to safely tweak and experiment with the system to achieve higher availability. We have been able to run millions of Paxos groups in a region with log replicas. Eventually, consistent reads can be served by any of the replicas. In case a leader fails, other replicas detect its failure and elect a new leader to minimize disruptions to the availability of consistent reads.

Research areas

Related content

US, CA, Santa Clara
The Data Intelligence team is a new function within Amazon Customer Service (CS). We own the end-to-end process of defining, building, implementing, and monitoring a comprehensive data strategy. We also develop and apply Generative Artificial Intelligence (GenAI), Machine Learning (ML), Ontology, and Natural Language Processing (NLP) to enhance customer service associate and customer experiences. As an Applied Scientist, you'll own the definition and implementation of customer-focused, AI-driven innovation in Amazon Customer Service globally, leveraging GenAI, ML, and/or NLP to transform complex business requirements and customer needs into innovative technology solutions. Your expertise will be key in shaping data-driven strategies and addressing complex data challenges. With your expertise in AI, text analysis, embeddings, language modeling, and generation, you'll design and develop scalable AI-powered technology solutions, prioritize initiatives, drive data-driven insights, and deliver business impact. This position will advance applied science best practices, leverage data and AI to drive customer experience improvements, and set new global standards for customer experience. This role requires you to work with a cross-functional team, including scientists, engineers, and product managers, to develop scalable and maintainable AI solutions for both structured and unstructured data. The ideal candidate has strong technical skills in AI techniques (e.g., automated reasoning, reasoning, planning, knowledge representation), excellent written documentation skills, and experience with big data technologies. Success in this role requires combining deep business knowledge with hands-on technical skills to solve customer problems and address complex technical challenges. Key job responsibilities - Develop innovative solutions to complex problems (e.g., Automated Reasoning for Trusted AI-Enabled Customer Service). - Apply technical expertise to implement novel algorithms and modeling solutions, in collaboration with other scientists and engineers. - Analyze data and define metrics to identify actionable insights and measure improvements in customer experience. - Communicate results and insights to both technical and non-technical audiences through written reports, presentations, and internal/external publications. - Collaborate with product management and engineering teams to integrate and optimize models in production systems. A day in the life A typical day as an Applied Scientist in the Data Intelligence team involves combining business expertise with hands-on problem-solving in ML and AI. The role encompasses tackling complex data initiatives, ensuring alignment with customer needs and business objectives, and translating business requirements into practical AI-driven solutions. Working collaboratively with cross-functional teams, this position involves designing and enhancing AI models, focusing on efficiency, precision, and scalability. Daily activities include ensuring data quality, monitoring model performance, and generating actionable insights from vast amounts of information. Each day presents opportunities to resolve complex technical challenges, advance important AI projects, and conceive innovative ways to leverage data in transforming the customer experience. About the team The Data Intelligence team is a new function within Amazon Customer Service. We develop and apply Generative Artificial Intelligence (GenAI), Machine Learning (ML), and Natural Language Processing (NLP) techniques to enhance customer service associate and customer experiences.
US, NY, New York
Application deadline: Applications will be accepted on an ongoing basis Applied Scientists in AWS Automated Reasoning develop and apply bleeding-edge formal methods, automated reasoning techniques, and neurosymbolic approaches to ensure the security, reliability, and correctness of Amazon/AWS services and customer applications. Our tools are called billions of times daily, powering the backbone of Amazon's products and services. We are changing the way computer systems are developed and operated, raising the bar for security, durability, availability, and quality. At Amazon, automated reasoning is central to maintaining customer trust and delivering delightful customer experiences. Application areas span cloud infrastructure verification, cryptographic assurance, AI safety, drone safety, and formal guarantees for generative AI systems. Our methods range from interactive theorem proving and constraint solving to neuro-inspired proof search This is a unique opportunity to get in early on a fast-growing segment of the business and help shape the technology, product, and business. You will have a chance to utilize your deep technical expertise within a fast-moving environment and make a large business and customer impact. Key job responsibilities • Design and implement algorithms and formal methods for automated reasoning, including constraint solving, model checking, static analysis, theorem proving, and program synthesis to verify the correctness, security, and reliability of computing systems. • Solve large or significantly complex problems that require deep knowledge and scientific innovation in your domain; own strategic problem solving and take the lead on design, implementation, and delivery of solutions with long-term quantifiable impact. • Develop new decision procedures, heuristics, and search strategies that improve the scalability and accuracy of verification tools; build and deploy production-grade automated reasoning systems at Amazon scale. • Explore and apply generative AI and machine learning techniques to enhance automated reasoning capabilities, including learning-based heuristics for search and optimization, neural approaches to symbolic reasoning, and methods for verifying the correctness of AI-generated code. • Develop automated reasoning techniques for generative AI and agentic coding systems, including methods for ensuring the safety and alignment of autonomous software agents and applying formal guarantees to large language model outputs. • Conduct original research snd publish findings in peer-reviewed venues. • Work with customer teams to understand the nature of their software and the properties they need to establish; identify tools and methods capable of addressing verification needs, including novel analysis capabilities. • Provide cross-organizational technical influence, increasing productivity and effectiveness by sharing deep knowledge and experience; collaborate with partner teams to translate verification capabilities into production systems. • Mentor scientists and engineers on formal methods, neurosymbolic techniques, and best practices for building reliable automated reasoning systems; assist in career development of others.
US, WA, Seattle
At Amazon Selection and Catalog Systems (ASCS), our mission is to power the online buying experience for customers worldwide so they can find, discover, and buy any product they want. We innovate on behalf of our customers to ensure uniqueness and consistency of product identity and to infer relationships between products in Amazon Catalog to drive the selection gateway for the search and browse experiences on the website. We're solving a fundamental AI challenge: establishing product identity and relationships at unprecedented scale. Using Generative AI, Visual Language Models (VLMs), and multimodal reasoning, we determine what makes each product unique and how products relate to one another across Amazon's catalog. The scale is staggering: billions of products, petabytes of multimodal data, millions of sellers, dozens of languages, and infinite product diversity—from electronics to groceries to digital content. The research challenges are immense. GenAI and VLMs hold transformative promise for catalog understanding, but we operate where traditional methods fail: ambiguous problem spaces, incomplete and noisy data, inherent uncertainty, reasoning across both images and textual data, and explaining decisions at scale. Establishing product identities and groupings requires sophisticated models that reason across text, images, and structured data—while maintaining accuracy and trust for high-stakes business decisions affecting millions of customers daily. Amazon's Item and Relationship Platform group is looking for an innovative and customer-focused applied scientist to help us make the world's best product catalog even better. In this role, you will partner with technology and business leaders to build new state-of-the-art algorithms, models, and services to infer product-to-product relationships that matter to our customers. You will pioneer advanced GenAI solutions that power next-generation agentic shopping experiences, working in a collaborative environment where you can experiment with massive data from the world's largest product catalog, tackle problems at the frontier of AI research, rapidly implement and deploy your algorithmic ideas at scale, across millions of customers. Key job responsibilities Key job responsibilities include: * Formulate novel research problems at the intersection of GenAI, multimodal learning, and large-scale information retrieval—translating ambiguous business challenges into tractable scientific frameworks * Design and implement leading models leveraging VLMs, foundation models, and agentic architectures to solve product identity, relationship inference, and catalog understanding at billion-product scale * Pioneer explainable AI methodologies that balance model performance with scalability requirements for production systems impacting millions of daily customer decisions * Own end-to-end ML pipelines from research ideation to production deployment—processing petabytes of multimodal data with rigorous evaluation frameworks * Define research roadmaps aligned with business priorities, balancing foundational research with incremental product improvements * Mentor peer scientists and engineers on advanced ML techniques, experimental design, and scientific rigor—building organizational capability in GenAI and multimodal AI * Represent the team in the broader science community—publishing findings, delivering tech talks, and staying at the forefront of GenAI, VLM, and agentic system research
US, CA, East Palo Alto
As part of the AWS Solutions organization, we have a vision to provide business applications, leveraging Amazon’s unique experience and expertise, that are used by millions of companies worldwide to manage day-to-day operations. We will accomplish this by accelerating our customers’ businesses through delivery of intuitive and differentiated technology solutions that solve enduring business challenges. We blend vision with curiosity and Amazon’s real-world experience to build opinionated, turnkey solutions. Where customers prefer to buy over build, we become their trusted partner with solutions that are no-brainers to buy and easy to use. Key job responsibilities Everyone on the team needs to be entrepreneurial, wear many hats and work in a highly collaborative environment that’s more startup than big company. We’ll need to tackle problems that span a variety of domains: computer vision, image recognition, machine learning, real-time and distributed systems. As an Applied Scientist, you will help solve a variety of technical challenges and mentor other scientists. You will tackle challenging, novel situations every day and given the size of this initiative, you’ll have the opportunity to work with multiple technical teams at Amazon in different locations. You should be comfortable with a degree of ambiguity that’s higher than most projects and relish the idea of solving problems that, frankly, haven’t been solved at scale before - anywhere. Along the way, we guarantee that you’ll learn a ton, have fun and make a positive impact on millions of people. A key focus of this role will be developing and implementing advanced visual reasoning systems that can understand complex spatial relationships and object interactions in real-time. You'll work on designing autonomous AI agents that can make intelligent decisions based on visual inputs, understand customer behavior patterns, and adapt to dynamic retail environments. This includes developing systems that can perform complex scene understanding, reason about object permanence, and predict customer intentions through visual cues. About the team Just Walk Out (JWO) is a new kind of store with no lines and no checkout—you just grab and go! Customers simply use the Amazon Go app to enter the store, take what they want from our selection of fresh, delicious meals and grocery essentials, and go! Our checkout-free shopping experience is made possible by our Just Walk Out Technology, which automatically detects when products are taken from or returned to the shelves and keeps track of them in a virtual cart. When you’re done shopping, you can just leave the store. Shortly after, we’ll charge your account and send you a receipt. Check it out at amazon.com/go. Designed and custom-built by Amazonians, our Just Walk Out Technology uses a variety of technologies including computer vision, sensor fusion, and advanced machine learning. Innovation is part of our DNA! Our goal is to be Earths’ most customer centric company and we are just getting started. We need people who want to join an ambitious program that continues to push the state of the art in computer vision, machine learning, distributed systems and hardware design.
US, WA, Seattle
Amazon Leo is a constellation of Low Earth Orbit satellites that will provide low-latency, high-speed broadband network connectivity to unserved and underserved communities around the world. We are looking for an Applied Scientist to join the founding cohort of the Engineering and R\&D team within Leo Infrastructure and IP Security. The team defends the manufacturing lines, launch sites, and global ground infrastructure behind the constellation from the most sophisticated threat actors on the planet. The data is unlike anything you have worked with: badge and door-access events, asset movement, network telemetry, and industrial control signals from factories, ground stations, and launch facilities, all of which must be modeled, baselined, and defended. You will build the statistical and behavioral models that separate threat actor behavior from the noise of a global operation, and the privacy-preserving data representations that let detection science scale without exposing sensitive data. This is an R\&D role with a production mandate: every model you build becomes part of the system Leo's security teams use to protect the constellation. #### Export Control Requirement Due to applicable export control laws and regulations, candidates must be a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum. Key job responsibilities - Build behavioral and statistical models that baseline normal activity across heterogeneous security telemetry, including specialized industrial-control and factory-floor data sources, so detections extend to new environments with low false-positive rates. - Design privacy-preserving representations of sensitive security data, and verify that models and detections tuned against them remain accurate against the real data — turning data-protection guarantees into measurable, provable properties rather than assertions. - Define the methodology and own the analysis for difficult, loosely defined problems: gather complex data across domains, select the right techniques from a range of data science methods, and justify your approach with evidence. - Develop the metrics and evaluation frameworks that measure model and detection performance against threat actor behavior before a model is trusted in production. - Contribute to the team's neurosymbolic reasoning platform, adapting state-of-the-art techniques from the literature and shipping components at production quality. - Document your work with the rigor of a peer-reviewed publication, and communicate results clearly to both scientific and security-operations audiences. A day in the life You will move between analysis and production in the same week: profiling a telemetry source the team has never modeled, establishing what normal looks like for it, and shipping the baseline as a component detections build on. Security engineers on your team translate threat intelligence into the adversary behaviors that matter; you build and tune the models that detect those behaviors and evaluate model performance against them. You might spend a morning chasing a false-positive pattern to its statistical root cause, and the afternoon verifying that an anonymized dataset still preserves the signal a detection depends on. You will work semi-autonomously with guidance from senior scientists, backtest candidate detections against retained telemetry, and deliver scientific artifacts that ship. About the team Leo Infrastructure and IP Security protects the people, facilities, hardware, and supply chain behind a global satellite constellation. The Engineering and R\&D team within this organization builds the platforms and tooling the security pillar teams operate on, moving security operations from manual triage to correlation-based detection, automated response, and agentic AI. The team is composed of applied scientists, software engineers, and security engineers working across physical and digital security domains. #### Inclusive Team Culture In Amazon Security, it's in our nature to learn and be curious. Ongoing DEI events and learning experiences inspire us to continue learning and to embrace our uniqueness. Addressing the toughest security challenges requires that we seek out and celebrate a diversity of ideas, perspectives, and voices. #### Training & Career Growth We're continuously raising our performance bar as we strive to become Earth's Best Employer. That's why you'll find endless knowledge-sharing, training, and other career-advancing resources here to help you develop into a better-rounded professional. #### Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why flexible work hours and arrangements are part of our culture. When we feel supported in the workplace and at home, there's nothing we can't achieve.
GB, London
Amazon’s Middle Mile Science group is looking for an Applied Scientist to build machine learning and optimization models to support pricing and revenue management of its external freight business. This includes the development of novel forecasting and dynamic pricing models, as well as the application of causal inference and artificial intelligence techniques, to improve marketplace services and execution for our customers. The Middle Mile Science group develops optimization and machine learning systems that power Amazon's freight transportation network, from network design and pricing to real-time load planning and capacity utilization. The scale of Amazon's fulfillment operations requires robust transportation networks that minimize cost while meeting all customer deadlines. Real-time execution depends on state-of-the-art optimization and artificial intelligence to coordinate thousands of operators and drivers. This includes shipper-facing and carrier-facing marketplace algorithms as well as network planning and optimization tools. Amazon often finds that existing techniques do not match our unique business needs,driving the innovation of new approaches and algorithms. As an Applied Scientist responsible for middle mile transportation, you will be working closely with different teams including business leaders and engineers to design and build scalable products operating across multiple transportation modes. You will create experiments and prototype implementations of new learning algorithms and prediction techniques. You will have exposure to top level leadership to present findings of your research. You will also work closely with other scientists and engineers to implement your models within our production system. You will implement solutions that are exemplary in terms of algorithm design, clarity, model structure, efficiency, and extensibility, and make decisions that affect the way we build and integrate algorithms across our product portfolio. About the team Our Middle Mile Marketplace Science team builds the algorithms for Amazon’s rapidly growing freight marketplace. Amazon contracts with 3P shippers and a network of independent carriers, using a mix of contract structures with varying service and risk profiles. Our work focuses on mechanisms and learning algorithms to optimize pricing and matching in this complex marketplace, and continually improve the experience for carriers and shippers. This is an area with many challenging problems and a huge business impact for Amazon!
US, CA, Pasadena
The Amazon Center for Quantum Computing in Pasadena, CA, is looking to hire a Fabrication R&D Scientist with experience in semiconductor process development who will aid in Amazon’s effort to bring cloud quantum computing services to its worldwide customer base. You will join a multi-disciplinary team of scientists, and hardware and software engineers working at the forefront of quantum computing. Through your work inside and outside of the cleanroom environment in the fabrication research and development group, you will solve problems related to developing next-generation quantum processors. Key job responsibilities Candidates must have a demonstrated background in sound scientific and engineering principles, and must have excellent data analysis, bias for action, problem solving, and communication skills, and be highly motivated and curious to research and learn new technical topics as needed. As a Fab R&D scientist you will be expected to work on new ideas and stay abreast of novel approaches in fabricating and packaging superconducting quantum processors. Working effectively within a team environment is critical. A day in the life The candidate will develop novel technologies using micro-/nano-fabrication techniques inside the cleanroom (independently or in collaboration with other scientists and engineers) for next-generation quantum computing. Outside the cleanroom, the candidate will plan experiments, analyze data, and conceive future innovations. About the team Candidates must have a demonstrated background in sound scientific and engineering principles, and must have excellent data analysis, bias for action, problem solving, and communication skills, and be highly motivated and curious to research and learn new technical topics as needed. As a Fab R&D scientist you will be expected to work on new ideas and stay abreast of novel approaches in fabricating and packaging superconducting quantum processors. Working effectively within a team environment is critical. Diverse Experiences Amazon values diverse experiences. Even if you do not meet all the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying. Work/Life Balance Our team puts a high value on work-life balance. It isn’t about how many hours you spend at home or at work; it’s about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives. Mentorship & Career Growth We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional. Export Control Requirement Due to applicable export control laws and regulations, candidates must be either a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum, or be able to obtain a US export license. If you are unsure if you meet these requirements, please apply and Amazon will review your application for eligibility. Key job responsibilities Responsibilities include developing and optimizing processes to fabricate high-coherence superconducting qubits; developing advanced 3DI interconnect and routing technologies for integrating superconducting quantum technologies; analyzing inline metrology and electrical test data; developing and maintaining integration documentation, design rules, and standard operating procedures; interacting with project leads to provide feedback that continuously improves different processes; staying updated with the latest advancements and industry trends in process integration and apply knowledge to improve processes and drive innovation providing technical guidance and support to junior colleagues, fostering a collaborative and knowledge-sharing work environment.
US, NY, New York
Sponsored Products and Brands at Amazon Ads is reimagining the advertising landscape through industry-leading generative AI technologies, revolutionizing how millions of customers discover products and engage with brands across Amazon.com and beyond. We are at the forefront of reinventing advertising experiences, bridging human creativity with artificial intelligence to transform every aspect of the advertising lifecycle, from ad creation and optimization to performance analysis and customer insights. We deliver billions of ad impressions and millions of clicks daily, and are breaking fresh ground to improve both the shopper and advertiser experience. We are a passionate group of innovators dedicated to developing responsible and intelligent AI technologies that balance the needs of advertisers, enhance the shopping experience, and strengthen the marketplace. The General Shopping Intelligence (GSI) team is a highly motivated, collaborative, and fun-loving group with a strong entrepreneurial spirit and bias for action. We provide advanced real-time machine learning services that connect shoppers with the right ads across all platforms and surfaces worldwide. Through deep understanding of both shoppers and products, we help shoppers discover new products they love, enable advertisers to reach their customers most efficiently, and help Amazon continuously innovate on behalf of all customers. We are seeking a motivated Applied Scientist who loves to innovate at the intersection of customer experience, deep learning, generative AI and high-scale machine learning systems. If you're energized by solving complex challenges and pushing the boundaries of what's possible with AI, join us in shaping the future of advertising. Key job responsibilities As an Applied Scientist, you will: * Conduct deep data analysis to derive insights to the business, and identify gaps and new opportunities * Develop scalable and effective machine-learning models and Generative AI solutions to solve business problems * Run regular A/B experiments, gather data, and perform statistical analysis * Work closely with software engineers to deliver end-to-end solutions into production * Improve the scalability, efficiency and automation of large-scale data analytics, model training, deployment and serving * Conduct research on new generative AI modeling to optimize all aspects of Sponsored Products and Brands business About the team We are pioneers in applying advanced machine learning and generative AI algorithms in Sponsored Products and Brands business. We empower every customer with a customized discovery experiences from back-end optimization (such as customized response prediction models) to front-end CX innovation (such as widgets), to help shoppers feel understood and shop efficiently on and off Amazon.
US, WA, Seattle
Amazon Leo is a constellation of Low Earth Orbit satellites that will provide low-latency, high-speed broadband network connectivity to unserved and underserved communities around the world. We are looking for an Applied Scientist to be a founding scientist on the Engineering and R\&D team within Leo Infrastructure and IP Security. The team defends the manufacturing lines, launch sites, and global ground infrastructure behind the constellation from the most sophisticated threat actors on the planet. These requirements create open scientific problems at the intersection of agentic AI, real-time stream processing, graph-based reasoning, and behavioral analytics. You will build the science behind a neurosymbolic reasoning platform and the models that detect the behavior of sophisticated threat actors. This is an R\&D role with a production mandate, where you define the problem rather than solve a pre-scoped one, and every model, detection, and agent workflow you build becomes the system Leo's security teams use to protect the constellation. #### Export Control Requirement Due to applicable export control laws and regulations, candidates must be a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum. Key job responsibilities - Design and implement scalable, production-grade neurosymbolic systems that integrate symbolic reasoning over graph-based knowledge representations with LLM agents to deliver reliable, verifiable security outcomes. - Design and run reinforcement learning and fine-tuning pipelines (GRPO, PPO, DPO) to optimize language models for security reasoning, triage, and detection-authoring tasks. - Build behavioral and statistical models that detect threat actor behavior, and design the evaluation frameworks that measure model performance against that behavior before trusting a model in production. - Design and build multi-agent systems that autonomously triage, enrich, and contain security events, including the constrained reasoning, safety guardrails, and validation mechanisms that make automated decisions trustworthy at scale. - Own the end-to-end science lifecycle, from research and experimentation through production deployment, defining the metrics that measure system performance and real-world security impact. - Advance the state of the art through publications at top-tier venues, patents, or open-source contributions, and shape the scientific agenda and research culture from day one. A day in the life You will move between research and production in the same week: framing an ambiguous security problem as a scientific question, prototyping an approach, and partnering with software engineers to ship it as a capability the platform runs continuously. Security engineers on your team translate threat intelligence into the adversary behaviors that matter; you build the models that detect those behaviors and evaluate model performance against them. You will obsess over the two latencies that define the platforms, the time from event to detection and the time from detection to containment action, and design agents and detections that drive both down. You will backtest candidate detections against retained telemetry, review evaluation results before a model or agent capability graduates to automated execution, and deliver scientific artifacts that ship. About the team Leo Infrastructure and IP Security protects the people, facilities, hardware, and supply chain behind a global satellite constellation. The Engineering and R\&D team within this organization builds the platforms and tooling the security pillar teams operate on, moving security operations from manual triage to correlation-based detection, automated response, and agentic AI. The team is composed of applied scientists, software engineers, and security engineers working across physical and digital security domains. #### Inclusive Team Culture In Amazon Security, it's in our nature to learn and be curious. Ongoing DEI events and learning experiences inspire us to continue learning and to embrace our uniqueness. Addressing the toughest security challenges requires that we seek out and celebrate a diversity of ideas, perspectives, and voices. #### Training & Career Growth We're continuously raising our performance bar as we strive to become Earth's Best Employer. That's why you'll find endless knowledge-sharing, training, and other career-advancing resources here to help you develop into a better-rounded professional. #### Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why flexible work hours and arrangements are part of our culture. When we feel supported in the workplace and at home, there's nothing we can't achieve.
US, WA, Seattle
Ever wonder how you can keep the world’s largest selection also the world’s safest and legally compliant selection? Then come join a team with the charter to monitor and classify the billions of items in the Amazon catalog to ensure compliance with various legal regulations. The Classification and Policy Platform team is looking for Sr. Applied Scientists to build technology to automatically monitor the billions of products on the Amazon platform. The software and processes built by this team are a critical component of building a catalog that our customers trust. You will have an opportunity to work with cutting edge machine learning algorithms on large datasets. You will need to build Amazon scale applications running on Amazon Cloud that both leverage and create new technologies to process large volumes of data that derive patterns and conclusions from the data. We are looking for highly motivated applied scientists and engineers interested in delivering the next level of innovation to product search for Amazon. As an Applied Scientist on the CPP team, you will be responsible for working across backend, client, business development, and data engineering teams to coordinate deep-dives, inform roadmaps, visualize metrics, and create predictive models to determine how we can best serve our customers. Responsibilities include: - Designing and implementing new features and machine learned models, including the application of state-of-art deep learning to solve search matching and ranking problems, including filtering, new content indexing, and apply document understanding - Conducting and coordinating process development leading to improved and streamlined processes for model development. Strong customer focus is essential - Working closely with Product Managers to expand depth of our product insights with data, create a variety of experiments, and determine the highest-impact projects to include in planning roadmaps - Providing technical and scientific guidance to your team members - Communicating effectively with senior management as well as with colleagues from science, engineering, and business backgrounds - Being a cultural leader that ensures teams are collecting, understanding, and using data to inform every decision that impacts our customers The successful candidate will have an established background in developing customer-facing experiences, a strong technical ability, a start-up mentality, excellent project management skills, and great communication skills. Amazon Science gives you insight into the company’s approach to customer-obsessed scientific innovation. Amazon fundamentally believes that scientific innovation is essential to being the most customer-centric company in the world. It’s the company’s ability to have an impact at scale that allows us to attract some of the brightest minds in artificial intelligence and related fields. Our scientists continue to publish, teach, and engage with the academic community, in addition to utilizing our working backwards method to enrich the way we live and work. Please visit https://www.amazon.science for more information.