Trainium x Reactor Rolling Forcing
The AWS Neuron Science team and Reactor collaborated to bring real-time video generation to Trainium.

A kernel-centric path to real-time video generation on Trainium

Using the Neuron Kernel Interface, a Reactor–AWS collaboration tackled the dynamic shapes, memory access patterns, and cache management that make real-time autoregressive diffusion hard—building techniques that generalize across models.

In 2018, Jürgen Schmidhuber — who in 1990 was the first person to propose world models as a machine learning concept — published a paper with David Ha. They wrote about “a predictive world model” which could “extract useful representations of space and time” and use that to train agents to drive, among other things. The growth in the capability and real-world applications of world models in the eight years since that seminal paper is staggering.

The emergence of massive training datasets, paired with variational autoencoders, and both diffusion and autoregressive transformers enable today’s world models to convert text, images, audio, and video inputs into latent space that is used to build and maintain breathtaking, immersive world models which can predict state changes and respond to actions. These models have vital and expanding roles in robotics, transport, climate modeling, video game development, scientific simulation, and design prototyping.

World models can generate immersive, explorable environments that respond to user input in real time.

The sharp increase in the power and potential of world models has been accompanied by an increased need for hardware and infrastructure capable of supporting them. That is a challenge that Bryce Schmidtchen, co-founder and chief technology officer of Reactor, is well acquainted with.

“It's about real time, it's about low latency, and it's about doing that as efficiently as you can at scale,” Schmidtchen observed. Reactor is a platform that allows developers, designers, and researchers to deploy, use, and scale real-time interactive AI models. “Efficiency means everything from how you schedule the inference on the given chip, in our case Trainium, to how you think about maximally bin packing every forward pass of the model.”

The challenge of efficiency is made even more acute by Reactor’s global presence. “We think about GPU clusters in terms of regions — we have hundreds all over the world. The latent that turns into a pixel that comes off the GPU needs to be able to get to the network without hopping around through Kubernetes. And you need to do this in a way where it doesn't matter if it's a well-supported cluster or bare metal in a closet.

“And then,” Schmidtchen continued, “you need to tie this to world-class networking and media streaming that supports different codecs, different resolutions, and allows that to connect to APIs and SDKs across different languages that can support all different types of applications.”

Trainium x Reactor Rolling Forcing
Reactor's platform provides APIs and SDKs across multiple languages and codecs, supporting real-time interactive AI applications — from gaming and robotics to film and research — on AWS.

The global reach of AWS already helps Reactor to achieve many of its goals. “Everything they have at the inference layer, the level of scale that we're able to achieve in different regions is what makes a platform of this scale possible and reliable,” Schmidtchen observed.

Now, a recent collaboration between Reactor and the Amazon Neuron Science team has forged a path to even greater efficiency for world models. This is a look at the optimization strategy those teams pursued, and how that laid the foundation for future models to achieve real-time capability on Trainium.

The rise of autoregressive diffusion models

“The Neuron Science team explores new techniques for generative AI model enablement optimization,” explained Jun Wu, a principal applied scientist on the Neuron team. “This might be a new model architecture or new algorithm for optimization or a new way to generate and optimize the code for running those models on Trainium. We identify opportunities to build a prototype and make it feasible to be ported to the production pipeline.”

The team spotted one such opportunity around the usage of diffusion models in video generation. “We noticed an evolution from generating shorter, fixed-length videos to infinite-length or dynamic-length videos,” Wu explained. “Generating high-quality video frames at those lengths gave rise to the usage of autoregressive diffusion models.”
Those types of models, which combine the sequential next-token prediction found in large language models with the iteration and refinement of diffusion models, are essential for users who want to generate video where they can navigate the generated environment.

“Autoregressive diffusion means a user keystroke can be absorbed as input to the model and, conditioned on the previously generated video frame, the model can correctly decide the next move or the next scene,” Wu observed. “Most of the interactive video generation models we are seeing today are using this.”

Wu and his team, Mason Fu and Lingfan Yu, both senior applied scientists, saw a chance to optimize how those models are deployed for real-time video generation on Trainium, and saw the opportunity to do this in collaboration with Reactor. “We had been working on real-time video and interactive video generation for over 18 months at that point,” Schmidtchen noted. The teams considered various video generation techniques and aligned on utilizing Rolling Forcing, citing both its ability to consistently generate high-quality 30-second videos and its relative size.

The hard part with these models isn't quality, it's that they have to run in real time. Unlike traditional video generation, which renders a full clip offline and returns it later, models like Rolling Forcing are streaming. Each frame is generated and immediately consumed, shown to a user or fed back in as the next input. That is how developers and customers actually use them, and a frame that arrives late breaks the experience, because generation has to stay ahead of the playback timeline. “High-quality generation diffusion models pose the challenge of a very long sequence, which requires a lot of memory consumption,” Wu explained. “Rolling Forcing is relatively small, but its sequence length is very large.” That combination of a small model, a very long sequence, and a hard frame-rate floor is what makes real time so demanding. Rolling Forcing's ability to generate 16 frames per second, the standard for video playback, meant it also met latency requirements. This is where Trainium adds value, bringing the performance and memory to sustain that long-sequence workload at speed and deliver real-time generation above 16 fps rather than merely producing good frames eventually. So the teams set about enabling Rolling Forcing, as a proxy for autoregressive diffusion video generation on Trainium.

A developer-friendly approach

The Neuron Science and Reactor teams adopted a kernel-centric, bottom-up methodology aimed at making life easier on developers. They focused on three challenges — dynamic shapes, unusual memory access patterns, and heavy cache management — that real-time video generation poses for generic compilers. Each of those challenges recurs on every forward pass, so what might be an insignificant delay on other workloads can compound into latency failures in real-time video generation.

For example, the challenges posed by dynamic shapes are partly rooted in the shifting nature of scenes in real-time video generation: Imagine a shot showing a pitcher, alone on the mound, that pans out to show the other players and then thousands of fans in the stands, all within seconds. Real-time video generation also involves frames which are generated and evicted in a continuous sliding-window process. “This means attention lengths vary across forward passes, and there are two distinct passes per window: denoising, then cache cleanup,” Wu said. “All of that makes static compilation hard.”

In addition to sliding KV cache copies, video generation workloads contain operations — rotary position embeddings (RoPE) and attention transposes — whose memory access patterns are workload-specific, making them challenging for any general-purpose compiler to fully optimize. For example, in video generation RoPE must contend with three axes (height, width, and time) rather than the single position it accounts for in LLMs. “The 3D rotary embedding interleaves odd and even elements along the innermost dimension, producing many tiny data transfers when compiled generically,” Wu explained.

Finally, because KV caching happens at every layer — each one reads and writes a rolling KV cache — high throughput is required for on-device copies. The need to read old cache contents and write new ones at every layer for every step acts as another significant drag on memory.

Using the Neuron Kernel Interface

To help solve for this additional complexity, the Neuron Science team turned to the Neuron Kernel Interface (NKI).
“NKI lets developers write compute kernels that run directly on NeuronCore hardware, with precise control over how data moves between memory and compute engines,” Wu explained. He noted that NKI gives developers a self-service path to fix hotspots directly, replacing specific bottleneck operations with hardware-tuned implementations where profiling shows that simply compiling models is insufficient.

“In the work we did, the 3D-RoPE kernel went from five seconds to 1.8 milliseconds, cache copies from 23 milliseconds to 1.9 milliseconds per layer, and attention transposes were eliminated entirely by fusing them into the attention kernel,” Wu noted. “For real-time workloads where every millisecond matters, that direct hardware access is what makes production-grade performance achievable.”

Additionally, NKI-Dev-Suite — an agent for generating NKI kernels — produced a working 3D-RoPE kernel on its first attempt. “The combined effect: the pipeline used 11 GB of high-bandwidth memory, while the standard eager-mode path ran out of memory,” Wu noted.

Hybrid sharding strategy

As established, video diffusion models produce token sequences far longer than text models under the real-time requirement. That makes the challenge of self-attention more acute. “Self-attention here operates on 23,400 query tokens attending to 32,760 context tokens, and accounts for about 70% of compute time,” Yu observed. “No single core handles this efficiently without distributing the work.”

To address this, the team used a hybrid sharding strategy entailing sequence parallelism (SP), or partitioning data sequentially, and tensor parallelism (TP), which shards tensors along a specific dimension to distribute computation across multiple devices. For certain non-self-attention parts of the module, the team utilized sequence parallelism. However, for the self-attention portions, only the hybrid approach sufficed.

“LLM attention is causal and unidirectional — each token attends only to previous tokens, the KV cache grows monotonically, and there's one attention type per layer with a single cache policy,” Wu said. Rolling Forcing, however, has two attention types per block: self-attention for spatiotemporal consistency across frames and cross-attention for text conditioning and bidirectional attention within the active window, since all frames are jointly refined from noise.
That, combined with a dual-policy KV cache (a sliding window for recent context plus a permanent attention sink for global context), two forward passes per window (denoising writes noisy KV entries, then a cache update pass overwrites them with clean values, ensuring future windows always attend to clean context), and 3D video token structure constrains how the sequence can be split across cores.

Yu explained that TP alone presents a math problem. “The WAN diffusion transformer model has 12 attention heads, but we have eight Neuron cores per chip, so it's not divisible. Using TP alone means you would have to pad, but padding wastes computation.”

SP alone, on the other hand, can break the 3D structures because, as Yu noted, “You cannot guarantee the sequence partition will be right at the boundary of a frame. Your data must be at least within the granularity of a frame, but if you partition on the frame boundary, those operations won't work.” The hybrid approach splits heads across 4 cores and sequences across 2, keeping both the math and the data layout correct. “The VAE decoder used spatial W-axis sharding, achieving a super-linear 8.25 times speedup,” Yu said.

Model structure changes

The Reactor and Neuron Science teams also optimized parts of the Rolling Forcing model code to run more efficiently on Trainium. “Basically, the model has two phases: one is diffusion, the other is the cache update,” Yu said. “Those are executed in two separate runs, but the issue is the cache update phase has significantly less computation, so if you execute it in a separate round, it has far less hardware utilization. This is because we also have to shard it, and so the high-level principle is that the less data you feed to the chip, the worse hardware utilization you have.”

The team optimized the model code so that the components each of those phases have in common were batched together.

“When we encounter components that are slightly different, we split again and then handle the different components separately. But for most of the pipeline, they are batched together,” Yu explained. “This is specifically useful for Trainium, because each instance has 16 chips and each chip has eight cores. And if you partition your computation across too many cores, each call will just have a small amount of partition computation, and that's underutilizing capacity.”

The result

After employing these, and other optimizations, Reactor and the Neuron Science teams were able to successfully generate a correct video on the first end-to-end run, utilizing Trainium to deliver real-time models. And, the teams emphasized, those results are generalizable.

Trainium x Reactor Rolling Forcing
Running Rolling Forcing on a single Trainium2 chip, the collaboration achieved real-time video generation — a result the teams say generalizes across autoregressive diffusion models.

“Rolling Forcing was the pipeline, it's a very small model,” said Yahav Biran, a principal solutions architect. “You can iterate quickly on it, but it's still a robust system end to end. It has all the complexity that you have in a robust system: the encoder, the DIT, the VAE, the decoder. So basically, if you take a more robust system, it is operating on the same building blocks.”

“We're building common techniques for models that employ autoregressive diffusion which also have a requirement for real-time interaction,“ Yu added. “We're not optimizing a single model only. We're building common techniques for supporting all models with the requirements of real-time streaming.”

The future

Schmidtchen said he is excited about the future this kind of work may enable. “In the not-so-far future, every pixel will be generated in real time, interactively,” he said. “Whole stories can be created by world models in real time: stories that react, that you can engage with, that can even change their entire landscape on the fly.”

He also noted that the work Amazon is doing, and has already done, will do a great deal to make those visions a reality.

“Trainium, is clearly showing a tremendous commitment from AWS and Amazon overall,” Schmidtchen noted. “There's a clearer roadmap of higher performance, better cost performance, more scale globally. In this future where you have real-time interactive AI that needs to be distributed at scale to consumers, physical AI, and more, Trainium is very well positioned to work very well at the inference layer—in terms of its parallelization and its memory and its software stack—and integrate nicely with AWS's global scale infrastructure. We are excited to continue working closely with Amazon as we explore the untapped potential of world models. This is just the beginning.”

Research areas

Related content

IL, Tel Aviv
Are you a scientist interested in pushing the state of the art in Information Retrieval, Large Language Models and Recommendation Systems? Are you interested in innovating on behalf of millions of customers, helping them accomplish their every day goals? Do you wish you had access to large datasets and tremendous computational resources? Do you want to join a team of capable scientist and engineers, building the future of e-commerce? Answer yes to any of these questions, and you will be a great fit for our team at Amazon. Our team is part of Amazon’s Personalization organization, a high-performing group that leverages Amazon’s expertise in machine learning, generative AI, large-scale data systems, and user experience design to deliver the best shopping experiences for our customers. Our team is building next-generation personalization systems powered by Large Language Models. We are tackling novel research challenges to help customers discover products they'll love - at Amazon scale and latency requirements. We are a team uniquely placed within Amazon, to have a direct window of opportunity to influence how customers will think about their shopping journey in the future. As an Applied Science Manager, you will lead a team of scientists working at the frontier of LLM-based personalization. You will set the technical vision, drive the research agenda, and ensure your team delivers production-ready solutions. You will hire, mentor, and develop world-class scientists while fostering a culture of innovation and scientific rigor. You will partner closely with engineering and product teams to translate ambitious research into customer-facing impact, and represent your team's work to senior leadership. Please visit https://www.amazon.science for more information.
BR, SP, Sao Paulo
Do you feel the challenge and the adrenaline kick when a huge data-set stares you in the face and you know that somewhere inside are hidden very important business insights that can fundamentally alter the way top business leaders think and act? Do you enjoy presenting strong data backed insights to business leaders; insights that can topple their long held beliefs and compel them to change their direction completely? If yes, then you are the one we are looking for. We are looking to invite passionate leaders, with expertise in generate power business insights from very large datasets, on a journey where the primary aim would be to enable needle moving business impacts through statistical analysis. We are looking for leaders who can envision the design and development of analytical infrastructure which can support strategic and tactical decision-making. Those who join this high visibility team would have to navigate through significant ambiguity in defining business problems and converting them to analytical problems. This role requires additional exposure and experience to Machine Learning. Key job responsibilities Use machine learning and analytical techniques to create scalable solutions for business problems • Analyze and extract relevant information from large amounts of Amazon’s historical business data to help automate and optimize key processes • Design, development, evaluate and deploy innovative and highly scalable ml models such as risk scorecards, income models, fraud models for predictive learning in credit risk applications • Research and implement novel machine learning and statistical approaches • Work closely with software engineering teams to drive real-time model implementations and new feature creations • Work closely with business owners and operations staff to optimize various business operations • Establish scalable, efficient, automated processes for large scale data analyses, model development, model validation and model implementation • Mentor other scientists and engineers in the use of ML techniques • Innovate with the latest GenAI technology to build highly automated solutions for efficient customer promotions • Design, develop and deploy end-to-end machine learning solutions in the Amazon production environment to delight Amazon customers • Collaborate with cross-functional teams to develop comprehensive ML/statistical models that can scale to millions of customers to multiple countries Understand the credit risk data and evaluate the best ml model/ solution for dynamic business problems. About the team Brazil Payments is part of the International Emerging Stores Payments team and focuses on supporting the launch of new payment and financial products to our customers in Brazil.
US, NY, New York
Amazon Advertising drives billions of ad impressions and millions of clicks daily, powering discovery and sales for advertisers across Amazon's Retail and Marketplace businesses. The Ads Marketing Decision Science team sits at the intersection of data science and marketing strategy. We build intelligent, data-driven systems that analyze advertiser behavior at large scale to deliver the right guidance to the right advertiser at the right time. Our work spans behavioral modeling, content intelligence, automated decision systems, and GenAI applications, enabling personalized marketing experiences that help advertisers make smarter advertising decisions and grow their business on Amazon. We are looking for a Data Scientist who brings strong fundamentals in machine learning, causal inference, and statistical modeling to solve real advertiser problems. You will build predictive models, design experiments, develop segmentation frameworks, and leverage GenAI capabilities where applicable, taking solutions end-to-end from proof-of-concept to production at scale. You will partner closely with scientists, engineers, and product managers on a daily basis to prototype rapidly, ensure data integrity in production systems, and deliver measurable advertiser impact. If you are passionate about solving real-world problems with next level science, come join us as we innovate and make history. Key job responsibilities • Define and execute data science solutions end-to-end, from problem framing through production deployment. • Build machine learning models (classification, regression, clustering, ranking) for advertiser segmentation, propensity modeling, and recommendations. • Apply causal inference and experimentation methods (A/B testing, difference-in-differences, propensity score matching) to measure the impact of marketing interventions. • Analyze large-scale advertiser behavioral data to identify trends, surface growth opportunities, and support optimal decision making. • Collaborate with colleagues across science and engineering disciplines for fast turnaround proof-of-concept prototyping at scale. • Establish and drive data hygiene best practices to ensure coherence and integrity of data feeding into production ML/AI solutions. • Leverage GenAI and LLM capabilities to enhance science products where applicable A day in the life You will solve real-world problems by analyzing large volumes of advertiser data, building predictive models, designing experiments, and measuring business impact. You will prototype rapidly, validate ideas with data, and partner with engineers to productize and scale successful solutions. You will collaborate daily with scientists, engineers, and product managers across the advertising organization, working in a cross-functional, fast-paced environment where data drives decisions and helps advertisers grow. About the team We are a team of Applied Scientists, Research Scientists, Data Scientists, and Business Intelligence Engineers with deep expertise in ML, NLP, Gen-AI, RL, and causal inference, from a diverse range of backgrounds. We partner closely with strong engineers, product managers, and sales leaders who bring ads-industry depth and experience building scalable modeling and software solutions.
ES, Madrid
We're building the intelligence behind how customers discover, trust and enjoy products on Amazon. We're solving complex catalogue quality challenges with machine learning, enhancing product discovery through computer vision and multimodal AI and pioneering agentic systems that autonomously navigate and stress-test the Amazon shopping experience to surface insights at scale. We're looking for PhD students across multiple research domains to invent, design, and implement state of the art solutions for never before solved problems. Your work here won't just stay in a notebook, it ships to production and reaches customers worldwide. Check out the details below including the job responsibilities, team details, and basic qualifications before submitting your application. You can find more information about the Amazon Science community as well as interview preparation tips via the links below; - https://www.amazon.science/ - https://amazon.jobs/content/en/career-programs/university/science - https://amazon.jobs/content/en/how-we-hire/university-roles/applied-science Key job responsibilities As an Applied Science Intern, you will own the design and development of end-to-end systems. You'll have the opportunity to write technical white papers, create roadmaps and drive production level projects that will support Amazon Science. You will work closely with Amazon scientists and other science interns to develop solutions and deploy them into production. You will have the opportunity to design new algorithms, models, or other technical solutions whilst experiencing Amazon's customer focused culture. The ideal intern should have the ability to work with diverse groups of people and cross-functional teams to solve complex business problems. A day in the life You'll spend your first weeks scoping your project with your mentor, then own the research and implementation end-to-end. Your work could involve developing machine learning and data analysis solutions that detect and resolve catalogue quality issues at massive scale, building computer vision and multimodal learning models that transform how customers discover and interact with products, engineering universal ML systems that make shopping on Amazon easier and more visually delightful, or creating autonomous agentic shoppers that tirelessly navigate the Amazon website to provide feedback and actionable insights — depending on the team you're matched with. Many interns publish at top-tier conferences or see their work deployed to production before the internship ends. Interns may also be considered for a return offer at the end of their internship, subject to performance evaluation and headcount availability. Further benefits of an Amazon Science internship include; - All of our internships offer a competitive salary - Interns are paired with an experienced manager and mentor(s) - Interns get invited to different intern program or office events - Interns can build their professional and personal network with other Amazon Scientists - Interns can potentially publish work at top tier conferences About the team We're hiring interns for multiple teams in Spain including but not limited to; •Tamale - Using Machine Learning and Data analysis solutions to solve complex catalogue quality problems • NintAI- Developing AI solutions, focusing on computer vision and multimodal learning to enhance how customers discover and interact with products • Home Innovation tech- Building universal, state of the art Machine Learning technology that makes shopping on Amazon easier and more visually delightful for our customers • EU Intech- Pioneers a population of agentic shoppers, autonomous AI agents, that tirelessly navigate and shop on the Amazon website, providing feedback and insights to improve the customer experience You'll submit a single application and we'll match you with science teams best aligned with your research interests. Applications are reviewed on a rolling basis, and your application stays active until we find a team match or confirm there are no matches available. Start dates are available throughout the year for durations of between 3–6 months. Please note, each team has different start date and duration preferences — your recruiter will confirm the preferences of the team you're matched with prior to interviewing. We offer science internships in multiple locations across the EMEA region and you can indicate your interest in all these locations by applying here (Austria, Estonia, France, Germany, Ireland, Israel, Italy, Jordan, Luxembourg, Netherlands, Poland, Romania, South Africa, Spain, Sweden, UAE, and UK). Please note we do not offer remote internships.
GB, London
We're building systems that turn research into real-world impact — ensuring flawless streaming for millions of Prime Video customers, shaping the future of AI-driven shopping with Rufus, advancing speech generation and conversational AI, optimizing large-scale infrastructure through intelligent observability, and transforming how people and jobs find each other. If this sounds interesting to you then we have a range of opportunities for you to explore. We're looking for PhD students across multiple research domains to invent, design, and implement state-of-the-art solutions for never-before-solved problems. Your work here won't just stay in a notebook — it ships to production and reaches customers worldwide. Check out the details below including the job responsibilities, team details, and basic qualifications before submitting your application. You can find more information about the Amazon Science community as well as interview preparation tips via the links below; - https://www.amazon.science/ - https://amazon.jobs/content/en/career-programs/university/science - https://amazon.jobs/content/en/how-we-hire/university-roles/applied-science Key job responsibilities As an Applied Science Intern, you will own the design and development of end-to-end systems. You'll have the opportunity to write technical white papers, create roadmaps and drive production level projects that will support Amazon Science. You will work closely with Amazon scientists and other science interns to develop solutions and deploy them into production. You will have the opportunity to design new algorithms, models, or other technical solutions whilst experiencing Amazon's customer focused culture. The ideal intern should have the ability to work with diverse groups of people and cross-functional teams to solve complex business problems. A day in the life You'll spend your first weeks scoping your project with your mentor, then own the research and implementation end-to-end. Your work could involve building computer vision models that detect quality issues across Prime Video's content library, developing multimodal AI solutions that power Amazon's shopping assistant Rufus, creating automated reasoning tools that verify distributed systems at scale, engineering intelligent observability frameworks for large-scale infrastructure, or advancing recommendation models that connect people with the right jobs — depending on the team you're matched with. Many interns publish at top-tier conferences or see their work deployed to production before the internship ends. Interns may also be considered for a return offer at the end of their internship, subject to performance evaluation and headcount availability. Further benefits of an Amazon Science internship include; - All of our internships offer a competitive salary - Interns are paired with an experienced manager and mentor(s) - Interns get invited to different intern program or office events - Interns can build their professional and personal network with other Amazon Scientists - Interns can potentially publish work at top tier conferences About the team We're hiring interns for multiple teams in the UK, including but not limited to; • Prime Video Video Quality Analysis – Develops AI and machine learning solutions using computer vision, audio processing, and generative AI to detect and prevent streaming quality issues across Prime Video's vast content library • Rufus Features Science UK – Shapes AI-driven shopping experiences at Amazon, working on projects from enabling Rufus to take actions on behalf of customers to generating multimodal answers combining text, image, audio, and video. • READI – Observability, Triage & Peak Readiness – Builds intelligent log analytics and automated performance frameworks that transform system telemetry into actionable insights for large-scale distributed systems. • Automated Reasoning Group – Ensures program and systems correctness through deductive proof, model checking, formal verification, and runtime conformance monitoring for large-scale distributed systems. You'll submit a single application and we'll match you with science teams best aligned with your research interests. Applications are reviewed on a rolling basis, and your application stays active until we find a team match or confirm there are no matches available. Start dates are available throughout the year for durations of between 3–6 months. Please note, each team has different start date and duration preferences — your recruiter will confirm the preferences of the team you're matched with prior to interviewing. We offer science internships in multiple locations across the EMEA region and you can indicate your interest in all these locations by applying here (Austria, Estonia, France, Germany, Ireland, Israel, Italy, Jordan, Luxembourg, Netherlands, Poland, Romania, South Africa, Spain, Sweden, UAE, and UK). Please note we do not offer remote internships.
IL, Tel Aviv
We're building systems that turn research into real-world impact — powering sports experiences for millions of Prime Video customers, pioneering multimodal document intelligence, building autonomous AI agents that reason, plan, and act, and transforming how customers discover products they love. If this sounds interesting to you then we have a range of opportunities for you to explore. We're looking for PhD students across multiple research domains to invent, design, and implement state-of-the-art solutions for never-before-solved problems. Your work here won't just stay in a notebook — it ships to production and reaches customers worldwide. Check out the details below including the job responsibilities, team details, and basic qualifications before submitting your application. You can find more information about the Amazon Science community as well as interview preparation tips via the links below; - https://www.amazon.science/ - https://amazon.jobs/content/en/career-programs/university/science - https://amazon.jobs/content/en/how-we-hire/university-roles/applied-science Key job responsibilities As an Applied Science Intern, you will own the design and development of end-to-end systems. You'll have the opportunity to write technical white papers, create roadmaps and drive production level projects that will support Amazon Science. You will work closely with Amazon scientists and other science interns to develop solutions and deploy them into production. You will have the opportunity to design new algorithms, models, or other technical solutions whilst experiencing Amazon's customer focused culture. The ideal intern should have the ability to work with diverse groups of people and cross-functional teams to solve complex business problems. A day in the life You'll spend your first weeks scoping your project with your mentor, then own the research and implementation end-to-end. Your work could involve building computer vision models that power live sports experiences for Prime Video, developing multimodal GenAI solutions for AWS document intelligence, creating agentic AI systems that reason and act autonomously, or advancing recommendation models that transform how customers discover content — depending on the team you're matched with. Many interns publish at top-tier conferences or see their work deployed to production before the internship ends. Interns may also be considered for a return offer at the end of their internship, subject to performance evaluation and headcount availability. Further benefits of an Amazon Science internship include; - All of our internships offer a competitive salary - Interns are paired with an experienced manager and mentor(s) - Interns get invited to different intern program or office events - Interns can build their professional and personal network with other Amazon Scientists - Interns can potentially publish work at top tier conferences About the team We're hiring interns for multiple teams in Israel, including but not limited to; • Prime Video Sports — Build innovative sports experiences for Prime Video, spanning computer vision, 3D simulation, and personalized content recommendations. • Personalization — Leverage LLMs, NLP, and recommender systems to match customers with products that align with their passions and shopping preferences. • Agentic AI — Build next-generation agentic AI systems that automate real-world knowledge work, spanning retrieval-augmented generation, multi-agent workflows, long-term memory and personalization, knowledge-graph construction, and agent evaluation. • DS3 Textract — Develop multimodal generative AI algorithms that pioneer state-of-the-art document understanding solutions impacting millions of customers. You'll submit a single application and we'll match you with science teams best aligned with your research interests. Applications are reviewed on a rolling basis, and your application stays active until we find a team match or confirm there are no matches available. Start dates are available throughout the year. Some teams offer full-time internships (3–6 months) while others offer part-time positions (50–60%, 8–12 months) — your recruiter will confirm the format during team matching. We offer science internships in multiple locations across the EMEA region and you can indicate your interest in all these locations by applying here (Austria, Estonia, France, Germany, Ireland, Israel, Italy, Jordan, Luxembourg, Netherlands, Poland, Romania, South Africa, Spain, Sweden, UAE, and UK). Please note we do not offer remote internships.
US, CA, Sunnyvale
We are seeking an Applied Scientist to focus on Robot Navigation. In this role, you'll research and develop advanced navigation systems that enable robots to move reliably and safely through complex, dynamic environments. You'll work across a broad spectrum of navigation approaches—from classical methods to learning-based techniques and foundation models—to build robust solutions for autonomous robot navigation. Key job responsibilities - Develop and implement robust navigation systems that enable reliable autonomous operation in complex, dynamic indoor environments with static and dynamic obstacles - Build simulation-based and on-device evaluation frameworks with comprehensive benchmarks and metrics for systematic comparison of navigation methods - Conduct sim-to-real transfer experiments, analyzing performance gaps and developing techniques to ensure reliable real-world navigation performance - Collaborate with world model, manipulation, and other teams to ensure seamless integration of navigation capabilities into the full robot system - Stay current with the latest advances in robot navigation, spatial reasoning, and related fields, and apply relevant findings to improve system performance - Mentor fellow scientists and engineers while maintaining strong individual technical contributions About the team Fauna Robotics, an Amazon company, is building capable, safe, and genuinely delightful robots for everyday life. Our goal is simple: make robots people actually want to live and interact with in everyday human spaces. We believe that future won’t arrive until building for robotics becomes far more accessible. Today, too much effort is spent reinventing the fundamentals. We’re changing that by developing tightly integrated hardware and software systems that make it faster, safer, and more intuitive to create real-world robotic products.
US, CA, Sunnyvale
Are you a passionate scientist who wants to build AI agents that make a real difference in people's lives? At Ring, our mission is to make neighborhoods safer, and we believe agentic AI will change how customers interact with their homes and communities. You'll invent agents that reason about real-world situations, take meaningful action, and keep customers in control, and then you'll see them reach millions of households. As an Applied Scientist, you'll work with talented peers to push the frontier of agentic AI. You'll build agents that turn the multimodal signals captured by Ring devices into understanding and action. You'll tackle open problems in planning, tool use, learning from feedback, and reliability, taking ideas from research all the way to deployment at scale. You'll collaborate with teams across Amazon to advance the science of customer experiences through highly optimized, integrated hardware and software platforms. Key job responsibilities - Design, develop, and deploy LLM-based agents that plan and carry out multi-step tasks for customers using tools, services, and device data. - Advance the state of the art in agent capabilities such as planning, tool use, memory, and learning from feedback (e.g., RL and agent fine-tuning), and publish where appropriate. - Partner with engineering, product, and science teams to turn research into agent-driven experiences that help keep homes and neighborhoods safer.
US, WA, Seattle
We are seeking a Senior Applied Scientist to join our team in developing pioneering AI research, Generative AI, Agentic AI, Large Language Models (LLMs), Diffusion and Flow Models, and other advanced Machine Learning and Deep Learning solutions for Amazon Selection and Catalog Systems, within the AI Lab Team. This role offers a unique opportunity to work on AI research and AI products that will shape the future of online shopping experiences. Our team operates at the forefront of AI research and development, working on challenges that directly impact millions of customers worldwide. We push the boundaries of AI at both the foundational and application layers. As a Senior Applied Scientist, you will have the chance to experiment with LLMs and deep learning techniques, apply your research to solve real-world problems at an unprecedented scale, and collaborate with experienced scientists to contribute to Amazon's scientific innovation. Join us in redefining the future of shopping. Your work will directly influence how customers interact with the world's largest online store. Key job responsibilities - Design and implement novel AI solutions for Amazon catalog of products - Develop and train state-of-the-art LLMs, Diffusion Models, and other Generative AI models - Build and deploy autonomous AI Agents in Amazon production ecosystem - Scale AI models to handle billions of diverse products across multiple languages and geographies - Conduct research in areas such as Autonomous AI Agents, Generative AI, Language Modeling, Multi-modality Computer Vision, Diffusion Models, Reinforcement Learning - Collaborate with cross-functional teams to integrate AI models into Amazon's production ecosystem - Contribute to the scientific community through publications and conference presentations
US, NY, New York
We are seeking a Senior Applied Scientist to lead research and development of novel security validation and monitoring techniques for AI systems at scale. You will own and contribute to four critical workstreams: 1. Real-Time Agent Monitoring Design and implement scientific approaches for continuous behavioral analysis of AI agents in production—detecting anomalous actions, prompt injection exploitation, and policy violations in real time. 2. MCP Server Validation Develop novel validation frameworks to assess the security posture of Model Context Protocol (MCP) servers, including input sanitization verification, tool-use authorization boundaries, and data exfiltration detection. 3. AI-Enabled Application Validation Invent and deliver scalable methodologies for security testing of AI-enabled applications, including adversarial robustness evaluation, safety guardrail bypass detection, and trust boundary verification. 4. AI Asset Discovery & Inventory Research and build scalable techniques to automatically discover, identify, and catalog all AI-enabled applications and services across the company—maintaining a comprehensive, continuously updated database of AI assets. Key job responsibilities Invent • Identify and frame new research challenges in AI security where problems are ill-defined and require novel scientific paradigms at the product level. • Drive the team's scientific agenda for agent monitoring, validation research, and AI asset discovery; propose new initiatives and secure leadership buy-in. • Publish research results at peer-reviewed internal and external venues (e.g., USENIX Security, IEEE S&P, NeurIPS, ICML security workshops) when appropriate. • Articulate key scientific challenges of current and future AI security threats and present interventions to address them. • Make trade-offs between short-term tactical security needs and long-term research investments. Implement • Lead the design, implementation, and successful delivery of scientifically complex security solutions into production—both brand new systems and evolutions of existing ones. • Write significant portions of critical-path code for detection models, validation engines, and asset discovery / classification systems. • Independently assess and select appropriate technologies (e.g., streaming inference frameworks, graph-based anomaly detection, NLP-based service classification, code/traffic analysis for AI fingerprinting) for production systems. • Drive adoption of best practices in scientific methodology and software engineering across the team; provide insightful peer reviews of code, design, and architecture artifacts. • Deliver solutions that are inventive, maintainable, scalable, and extensible. Influence • Autonomously drive discussions with security engineers, product managers, and scientist peers across multiple teams. • Build consensus on larger cross-team security initiatives and factor complex efforts into independent workstreams. • Proactively identify and resolve endemic problems, including areas where current security tooling limits innovation of partner teams. • Actively recruit, mentor, and develop other scientists; provide technical assessments for promotions. • Contribute to the broader internal and external scientific communities as a subject matter expert in AI security. About the team Diverse Experiences Amazon Security values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying. Why Amazon Security? At Amazon, security is central to maintaining customer trust and delivering delightful customer experiences. Our organization is responsible for creating and maintaining a high bar for security across all of Amazon’s products and services. We offer talented security professionals the chance to accelerate their careers with opportunities to build experience in a wide variety of areas including cloud, devices, retail, entertainment, healthcare, operations, and physical stores. Inclusive Team Culture In Amazon Security, it’s in our nature to learn and be curious. Ongoing DEI events and learning experiences inspire us to continue learning and to embrace our uniqueness. Addressing the toughest security challenges requires that we seek out and celebrate a diversity of ideas, perspectives, and voices. Training & Career Growth We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, training, and other career-advancing resources here to help you develop into a better-rounded professional. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why flexible work hours and arrangements are part of our culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve.