How some of AWS's most innovative customers are using computer vision technologies

From counting fish to identifying touchdowns, AWS customers are utilizing computer vision and pattern recognition technologies to improve business processes and customer experiences.

Computer vision, the automatic recognition and description of images and video, has applications that are far-reaching, from identifying defects in high speed assembly lines and its use in autonomous robots, to the analysis of medical images, and the identification of products and people in social media. This week, in line with the IEEE Computer Vision and Pattern Recognition (CVPR) conference, we’ve rounded up examples of how some of AWS's most innovative customers are utilizing computer vision and pattern recognition technologies to improve business processes and customer experiences. This includes approaches such as data scientists building custom vision models using Amazon SageMaker, and application developers using Amazon Rekognition and Amazon Textract to embed computer vision into their applications.

Advertising

REA Group image
REA Group has developed an image compliance system that automatically detects any noncompliance and notifies home sellers.
fstop123/Getty Images

In advertising and other online media, computer vision can automate content moderation. REA Group, a multinational digital advertising company specializing in property and real estate, provides search-based portals that enable property sellers to upload images of properties on the market to deliver a wide, searchable selection to their consumers. REA Group discovered that images uploaded to their portal often weren’t compliant with their usage terms. Some images included trademarks or contact details of the sellers, which created lead attribution challenges. They set up a dedicated team of individuals to manually review the images for unapproved content, but the large volume of daily uploads and the additional review process delayed the property listing time by several days. The REA team developed an image compliance system that automatically detects any noncompliance and notifies sellers. To augment their existing machine learning models, they're using Amazon Rekognition Text in Image, which detects and extracts text in images, enabling them to increase the accuracy of detecting noncompliance and reduce false positives by more than 56 percent. They added business rules that factored in a variety of predictions from their own models, and from Amazon Rekognition, to enable automated decision-making.

Agriculture

fin.png
Aquabyte's machine learning algorithms can estimate how much a fish weighs while still in the water.

Agriculture has also benefited from computer vision. Fish farming is one of the most efficient sources of protein, since a pound of feed equates to nearly a pound of protein. But the cold, dark waters of fish habitats make it nearly impossible to effectively manage these farms from the surface. Historically, fish farmers have had to randomly scoop fish out of the water to measure their weight and check for disease. Aquabyte’s machine learning solution reimagines this process by using underwater cameras that keep tabs on the fish and compare photos of them over time. The machine learning algorithms, running on Amazon SageMaker, can estimate how much each fish weighs while it’s still in the water. The system can also monitor the fish for sea lice, a parasite that is a major problem in salmon farms, and the subject of significant regulation in Norway, where the bulk of Aquabyte’s client base currently operates. Without a solution like Aquabyte, managing sea lice amounts to nearly a quarter of the cost of operating a salmon farm. Aquabyte’s cameras have counted 2 million sea lice to date, the result of billions of images being captured. The Aquabyte team has been working on methods that would allow farmers to track individual fish for growth-tracking and breeding purposes. In the future, machine learning might even help automate elements of the farms by intelligently distributing fish feed, for example.

Autonomous driving

grid.png
DeepMap is focused on solving the mapping and localization challenge for autonomous vehicles.

Industries like autonomous driving wouldn’t even be possible without the help of computer vision. Perhaps you think the world is already sufficiently mapped. With the advent of satellite images and Google Street View, it seems like every square inch of the globe is represented in data. But for autonomous vehicles, much of the world is uncharted territory. That’s because the maps designed for humans “can’t be consumed by robots,” says Tom Wang, the director of engineering at DeepMap, a Palo Alto startup focused on solving the mapping and localization challenge for autonomous vehicles. According to Wang, these new kinds of vehicles need higher precision maps with richer semantics, things like the traffic signals, a lot of different traffic signs, driving boundaries, and connecting lanes. For DeepMap computer vision is critical. DeepMap needs to run a vast volume of image detections to automatically generate a comprehensive list of map features and detect dynamic road changes. Using Amazon SageMaker, DeepMap updates training models within a day and runs image detection on tens of millions of images on a daily basis to keep up with ever-changing conditions.

Education

Certipass, a UNI ISO standards accredited body for the certification of digital skills
Certipass was able to build their solution in under 30 days, enabling all their testing centers to test candidates online during the COVID-19 pandemic.
fizkes/Getty Images/iStockphoto

In the wake of the COVID-19 pandemic, many educational institutions needed to quickly pivot to the online proctoring of exams, leading to a need for new ways to verify identification. Certipass, a UNI ISO standards accredited body for the certification of digital skills, is the primary provider of the international digital competency certification –European Informatics Passport (EIPASS).

Since the EIPASS Certification is an international standard, Certipass has made it their mission to ensure maximum security, objectiveness, transparency, and fairness during the entire online evaluation process. Certipass used Amazon Rekognition for automated candidate identity verification during tests that are in line with e-Competence Framework for Information and Communication Technology (CEN) and The Digital Competence Framework for Citizens (Joint Research Centre). They were able to build the solution in under 30 days to enable all their testing centers to test candidates online during COVID-19.

Financial services

Aella Credit
Aella Credit provides easy access to credit in emerging markets using biometric, employer, and mobile phone data
Victor Karanja/Getty Images

In financial services, Aella Credit provides easy access to credit in emerging markets using biometric, employer, and mobile phone data. For those in emerging markets, identity verification and validation is one of the major challenges to accessing retail banking services. How can you know that people are who they say they are in communities that don't have proper identification systems? Aella Credit uses Amazon Rekognition to analyze images to verify a customer’s identity and give them access to financial and healthcare services with minimal friction. Amazon Rekognition helps to automate video and image analysis, with no machine learning expertise required. What would have taken days to verify someone’s identity manually, now happens in seconds. Customers can actually receive their loan in their account in less than five minutes, broadening access to credit.

Financial technology

To make sure users are getting the largest possible tax refund, Intuit incorporates machine learning throughout the TurboTax experience to help users file their taxes more efficiently. TurboTax uses machine learning to shorten the filing process, which takes an average of 13 hours.

Taxes image for AWS customer success story
TurboTax utilizes machine learning to shorten the filing process.
simpson33/Getty Images/iStockphoto

With Intuit’s computer vision capabilities supported by Amazon Textract, entering information from tax forms like W2s or 1099s takes seconds. Rather than a user having to enter form fields manually, the service scans pictures of the forms and digitizes them. Then, using contextual data from TurboTax’s existing database of tax codes and compliance forms, Amazon Textract verifies accuracy and identifies any anomalies or missing data for the user.

Healthcare

face.png
By combining the power of machine learning and computer vision, an interdisciplinary team of researchers at Duke University has created a faster, less expensive, more reliable, and more accessible system to screen children for autism spectrum disorder.

Machine learning plays a key role in many health-related realms - from providers and payers looking to expedite the care continuum to pharma and biotech researchers looking to reduce costs and speed up the drug discovery and disease detection process. Researchers at Duke Center for Autism and Brain Development are using machine learning to screen for autism spectrum disorder (ASD) in children. It’s critically important to diagnose ASD as early in a child’s development as possible — starting treatment for ASD at an age of 18 to 24 months can increase a child’s IQ by up to 17 points—in some cases moving them into the “average” child IQ range of 90-110 (or above it)—and, in turn, significantly improving their quality of life. Currently, the wait time for children to receive a diagnosis could be well after the child’s third birthday. By combining the power of machine learning and computer vision, powered by AWS, an interdisciplinary team of researchers at Duke University have created a faster, less expensive, more reliable, and more accessible system to screen children for ASD.

Media and entertainment

Computer vision technology is helping sports organizations like the National Football League (NFL) improve the game for fans. The NFL works with AWS to develop real-time, state-of-the-art cloud technology leveraging machine learning and artificial intelligence to increase the efficiency and pace of the game.

For example, deep learning and computer vision technologies are being explored to aid game officiating including real-time football tracking. Within days, AWS and NFL scientists were able to create custom training data sets of thousands of images extracted from NFL broadcast game footage using Amazon SageMaker Ground Truth.

NFL football
Deep learning and computer vision technologies are being explored by the NFL to aid game officiating, including real-time football tracking.
CREDIT: National Foottball League

Working with the Amazon ML Solutions Lab, Amazon SageMaker and GluonCV with MXNet were used to train and optimize several state-of-the-art deep learning-based object detection models such as Faster-RCNN and Yolov3, to accurately detect the football across video frames. This led to a first-of-its-kind football tracking model that performs well in a number of complex scenarios, such as when the ball is highly occluded or is partially visible in different camera angles.

The NFL also uses computer vision to more easily and quickly search through thousands of media assets. The NFL photo team, official photographers of the NFL, has millions of photos in archive and generates 500,000 photos each season. Manually, they were able to tag 50,000 images over 18 months. By using Amazon Rekognition custom face collection, text in image, object detection, and Custom Labels, an automated machine learning object detection service, they were able to apply detailed tags for players, teams, objects, action, jerseys, location, etc. to their entire photo collection in a fraction of time it took previously. This allowed them to make these photos searchable and usable to everyone in the company in ways that weren't possible before.

For Sportradar, the global provider of sports and intelligence for the betting and media industries providing data coverage from more than 200,000 events annually, advances in computer vision are an opportunity to expand the depth of sports data offered to customers and reduce the costs of data collection through automation.

Sports betting image for AWS customer success story
For Sportradar, advances in computer vision are an opportunity to expand the depth of sports data offered to customers and reduce the costs of data collection through automation.
scyther5/Getty Images/iStockphoto

Sportradar is investing in computer vision research both through internal development and external partnerships to build computer vision data collection capabilities with an initial focus on tennis, soccer and snooker. Working with the Amazon ML Solutions Lab, Sportradar is exploring the application of state-of-the-art deep learning models for automated match event detection in soccer, moving beyond player and ball localization to understanding the intent of the play in terms of what is happening in the game.

To bring this technology into production as it matures, Sportradar is leveraging AWS services including Amazon SageMaker, EKS, MSK, FSx and Amazon’s broad range of GPU and CPU compute instances for its computer vision processing pipeline. This infrastructure allows Sportradar's researchers to test and validate computer vision models at scale and bring models from the lab to production with minimal effort while delivering the low latency, reliability and scalability needed for live sports betting use cases.

You can find more ways that AWS customers are innovating with computer vision here. More information about Amazon's participation at CVPR is available here.

Related content

US, WA, Seattle
As part of the AWS Applied AI Solutions organization, we have a vision to provide business applications, leveraging Amazon's unique experience and expertise, that are used by millions of companies worldwide to manage day-to-day operations. We will accomplish this by accelerating our customers' businesses through delivery of intuitive and differentiated technology solutions that solve enduring business challenges. We blend vision with curiosity and Amazon's real-world experience to build opinionated, turnkey solutions. Where customers prefer to buy over build, we become their trusted partner with solutions that are no-brainers to buy and easy to use. Just Walk Out (JWO) is a new kind of store with no lines and no checkout—you just grab and go! Customers simply use the Amazon Go app to enter the store, take what they want from our selection of fresh, delicious meals and grocery essentials, and go! Our checkout-free shopping experience is made possible by our Just Walk Out Technology, which automatically detects when products are taken from or returned to the shelves and keeps track of them in a virtual cart. When you’re done shopping, you can just leave the store. Shortly after, we’ll charge your account and send you a receipt. Check it out at amazon.com/go. Designed and custom-built by Amazonians, our Just Walk Out Technology uses a variety of technologies including computer vision, sensor fusion, and advanced machine learning. Innovation is part of our DNA! Our goal is to be Earths’ most customer centric company and we are just getting started. We need people who want to join an ambitious program that continues to push the state of the art in computer vision, machine learning, distributed systems and hardware design. Key job responsibilities Everyone on the team needs to be entrepreneurial, wear many hats and work in a highly collaborative environment that’s more startup than big company. We’ll need to tackle problems that span a variety of domains: computer vision, image recognition, machine learning, real-time and distributed systems. As a Sr. Applied Scientist, you will help solve a variety of technical challenges and mentor other scientists. You will be the thought leader of the team. You will tackle challenging, novel situations every day and given the size of this initiative, you’ll have the opportunity to work with multiple technical teams at Amazon in different locations. You should be comfortable with a degree of ambiguity that’s higher than most projects and relish the idea of solving problems that, frankly, haven’t been solved at scale before - anywhere. Along the way, we guarantee that you’ll learn a ton, have fun and make a positive impact on millions of people. Develop a novel framework and advance the theory and practice of multi-object tracking, re-identification, person activity understanding, multi-modal foundation model, and generic video understanding Create innovative techniques for efficient visual processing that can scale to real-world applications Investigate approaches to reduce the computational and data requirements of visual AI systems About the team AWS Solutions As part of the AWS solutions organization, we have a vision to provide business applications, leveraging Amazon's unique experience and expertise, that are used by millions of companies worldwide to manage day-to-day operations. We will accomplish this by accelerating our customers' businesses through delivery of intuitive and differentiated technology solutions that solve enduring business challenges. we blend vision with curiosity and Amazon's real-world experience to build opinionated, turnkey solutions. Where customers prefer to buy over build, we become their trusted partner with solutions that are no-brainers to buy and easy to use. Diverse Experiences AWS values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying. Why AWS? Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses. Inclusive Team Culture AWS values curiosity and connection. Our employee-led and company-sponsored affinity groups promote inclusion and empower our people to take pride in what makes us unique. Our inclusion events foster stronger, more collaborative teams. Our continual innovation is fueled by the bold ideas, fresh perspectives, and passionate voices our teams bring to everything we do. Mentorship & Career Growth We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve.
US, CA, Sunnyvale
Prime Video is a first-stop entertainment destination offering customers a vast collection of premium programming in one app available across thousands of devices. Prime members can customize their viewing experience and find their favorite movies, series, documentaries, and live sports – including Amazon MGM Studios-produced series and movies; licensed fan favorites; and programming from Prime Video subscriptions such as Apple TV+, HBO Max, Peacock, Crunchyroll and MGM+. All customers, regardless of whether they have a Prime membership or not, can rent or buy titles via the Prime Video Store, and can enjoy even more content for free with ads. Are you interested in shaping the future of entertainment? Prime Video's technology teams are creating best-in-class digital video experience. As a Prime Video team member, you’ll have end-to-end ownership of the product, user experience, design, and technology required to deliver state-of-the-art experiences for our customers. You’ll get to work on projects that are fast-paced, challenging, and varied. You’ll also be able to experiment with new possibilities, take risks, and collaborate with remarkable people. We’ll look for you to bring your diverse perspectives, ideas, and skill-sets to make Prime Video even better for our customers. With global opportunities for talented technologists, you can decide where a career Prime Video Tech takes you! Prime Video is pioneering the use of Generative AI to empower the next generation of creatives. Our mission is to make world-class media creation accessible, scalable and efficient. We are seeking a Lead Applied Scientist who have demonstrated experience in spearheading & advancing state of the art models, particularly in Generative AI. Your role will be to deliver these innovations as production-ready systems at Amazon scale. Key job responsibilities As a Sr. Applied Scientist, you will lead end-to-end product journey, research and experimentation for this domain. You will be applying advanced machine learning techniques in Computer Vision, Multimedia Understanding and Generative AI. We're building the foundational technology stack, spanning diffusion and flow-matching models, 3D/4D scene and character generation, motion and camera control, and post-training alignment. Other responsibilities include: - Lead research and develop generative models for controllable synthesis across images, video, vector graphics, and multimedia - Innovate in advanced diffusion and flow-based methods (e.g., inverse flow matching, parameter efficient training, guided sampling, test-time adaptation) to improve efficiency, controllability, and scalability - Advance visual grounding, depth and 3D estimation, segmentation, and matting for integration into pre-visualization, compositing, VFX, and post-production pipelines - Design multimodal GenAI workflows including visual-language model tooling, structured prompt orchestration, agentic pipelines
US, CA, Sunnyvale
Prime Video is a first-stop entertainment destination offering customers a vast collection of premium programming in one app available across thousands of devices. Prime members can customize their viewing experience and find their favorite movies, series, documentaries, and live sports – including Amazon MGM Studios-produced series and movies; licensed fan favorites; and programming from Prime Video subscriptions such as Apple TV+, HBO Max, Peacock, Crunchyroll and MGM+. All customers, regardless of whether they have a Prime membership or not, can rent or buy titles via the Prime Video Store, and can enjoy even more content for free with ads. Are you interested in shaping the future of entertainment? Prime Video's technology teams are creating best-in-class digital video experience. As a Prime Video team member, you’ll have end-to-end ownership of the product, user experience, design, and technology required to deliver state-of-the-art experiences for our customers. You’ll get to work on projects that are fast-paced, challenging, and varied. You’ll also be able to experiment with new possibilities, take risks, and collaborate with remarkable people. We’ll look for you to bring your diverse perspectives, ideas, and skill-sets to make Prime Video even better for our customers. With global opportunities for talented technologists, you can decide where a career Prime Video Tech takes you! Prime Video is pioneering the use of Generative AI to empower the next generation of creatives. Our mission is to make world-class media creation accessible, scalable and efficient. We are seeking an Applied Scientist to advance the state of the art in Generative AI and to deliver these innovations as production-ready systems at Amazon scale. Your work will give creators unprecedented freedom and control while driving new efficiencies. Key job responsibilities As an Applied Scientist, you will have end-to-end ownership of the product, related research and experimentation. In addition, you will be applying advanced machine learning techniques in Computer Vision, Multimedia Understanding and Generative AI. We're building the foundational technology stack, spanning diffusion and flow-matching models, 3D/4D scene and character generation, motion and camera control, and post-training alignment. Other responsibilities include: - Research and develop generative models for controllable synthesis across images, video, vector graphics, and multimedia - Innovate in advanced diffusion and flow-based methods (e.g., inverse flow matching, parameter efficient training, guided sampling, test-time adaptation) to improve efficiency, controllability, and scalability - Advance visual grounding, depth and 3D estimation, segmentation, and matting for integration into pre-visualization, compositing, VFX, and post-production pipelines - Design multimodal GenAI workflows including visual-language model tooling, structured prompt orchestration, agentic pipelines
IN, HR, Gurugram
Lead ML teams building large-scale forecasting and optimization systems that power Amazon’s global transportation network and directly impact customer experience and cost. As an Sr Applied Scientist, you will set scientific direction, mentor applied scientists, and partner with engineering and product leaders to deliver production-grade ML solutions at massive scale. Key job responsibilities 1. Lead and grow a high-performing team of Applied Scientists, providing technical guidance, mentorship, and career development. 2. Define and own the scientific vision and roadmap for ML solutions powering large-scale transportation planning and execution. 3. Guide model and system design across a range of techniques, including tree-based models, deep learning (LSTMs, transformers), LLMs, and reinforcement learning. 4. Ensure models are production-ready, scalable, and robust through close partnership with stakeholders. Partner with Product, Operations, and Engineering leaders to enable proactive decision-making and corrective actions. 5. Own end-to-end business metrics, directly influencing customer experience, cost optimization, and network reliability. 6. Help contribute to the broader ML community through publications, conference submissions, and internal knowledge sharing. A day in the life Your day includes reviewing model performance and business metrics, guiding technical design and experimentation, mentoring scientists, and driving roadmap execution. You’ll balance near-term delivery with long-term innovation while ensuring solutions are robust, interpretable, and scalable. Ultimately, your work helps improve delivery reliability, reduce costs, and enhance the customer experience at massive scale.
US, NY, New York
We are seeking an Research Scientist to lead the development of evaluation frameworks and data collection protocols for robotic capabilities. In this role, you will focus on designing how we measure, stress-test, and improve robot behavior across a wide range of real-world tasks. Your work will play a critical role in shaping how policies are validated and how high-quality datasets are generated to accelerate system performance. You will operate at the intersection of robotics, machine learning, and human-in-the-loop systems, building the infrastructure and methodologies that connect teleoperation, evaluation, and learning. This includes developing evaluation policies, defining task structures, and contributing to operator-facing interfaces that enable scalable and reliable data collection. The ideal candidate is highly experimental, systems-oriented, and comfortable working across software, robotics, and data pipelines, with a strong focus on turning ambiguous capability goals into measurable and actionable evaluation systems. Key job responsibilities - Design and implement evaluation frameworks to measure robot capabilities across structured tasks, edge cases, and real-world scenarios - Develop task definitions, success criteria, and benchmarking methodologies that enable consistent and reproducible evaluation of policies - Create and refine data collection protocols that generate high-quality, task-relevant datasets aligned with model development needs - Build and iterate on teleoperation workflows and operator interfaces to support efficient, reliable, and scalable data collection - Analyze evaluation results and collected data to identify performance gaps, failure modes, and opportunities for targeted data collection - Collaborate with engineering teams to integrate evaluation tooling, logging systems, and data pipelines into the broader robotics stack - Stay current with advances in robotics, evaluation methodologies, and human-in-the-loop learning to continuously improve internal approaches - Lead technical projects from conception through production deployment - Mentor junior scientists and engineers
US, CA, Sunnyvale
Prime Video is a first-stop entertainment destination offering customers a vast collection of premium programming in one app available across thousands of devices. Prime members can customize their viewing experience and find their favorite movies, series, documentaries, and live sports – including Amazon MGM Studios-produced series and movies; licensed fan favorites; and programming from Prime Video add-on subscriptions such as Apple TV+, Max, Crunchyroll and MGM+. All customers, regardless of whether they have a Prime membership or not, can rent or buy titles via the Prime Video Store, and can enjoy even more content for free with ads. Are you interested in shaping the future of entertainment? Prime Video's technology teams are creating best-in-class digital video experience. As a Prime Video technologist, you’ll have end-to-end ownership of the product, user experience, design, and technology required to deliver state-of-the-art experiences for our customers. You’ll get to work on projects that are fast-paced, challenging, and varied. You’ll also be able to experiment with new possibilities, take risks, and collaborate with remarkable people. We’ll look for you to bring your diverse perspectives, ideas, and skill-sets to make Prime Video even better for our customers. With global opportunities for talented technologists, you can decide where a career Prime Video Tech takes you! We are looking for a self-motivated, passionate and resourceful Applied Scientist to bring diverse perspectives, ideas, and skill-sets to make Prime Video even better for our customers. You will spend your time as a hands-on machine learning practitioner and a research leader. You will play a key role on the team, building and guiding machine learning models from the ground up. At the end of the day, you will have the reward of seeing your contributions benefit millions of Amazon.com customers worldwide. Key job responsibilities Develop foundation models for content understanding using state-of-the-art deep learning and multimodal learning techniques to analyze video, audio, and text. Build time sequence foundation models to understand and predict customer behavior patterns and viewing trajectories. Work closely with engineers and product managers to design, implement and launch solutions end-to-end across various Prime Video experiences. Design and conduct offline and online (A/B) experiments to evaluate proposed solutions based on in-depth data analyses. Effectively communicate technical and non-technical ideas with teammates and stakeholders. Stay up-to-date with advancements and the latest modeling techniques in foundation models, multimodal learning, and time series analysis. Publish your research findings in top conferences and journals. About the team Prime Video Recommendation Science team owns science solution to power recommendation and personalization experience on various Prime Video surfaces and devices. We work closely with the engineering teams to launch our solutions in production.
US, NY, New York
We are seeking an Applied Scientist to develop and optimize Visual Inertial Odometry (VIO) and sensor fusion systems for our intelligent robots. In this role, you will design, implement, and deploy state estimation and tracking algorithms that enable robots to understand their position and motion in real time, even in challenging and dynamic environments. You will own the full pipeline from algorithm development through embedded deployment, ensuring that perception systems run efficiently on resource-constrained robotic hardware. You will also leverage modern machine learning approaches to push the boundaries of classical perception methods, combining learned representations with geometric techniques to achieve robust, real-time performance. This is a deeply hands-on role. You will work directly with sensors, hardware, and real-world data, while prototyping, testing, and iterating in physical environments. The ideal candidate has strong foundations in VIO and sensor fusion, practical experience optimizing algorithms for embedded platforms, and familiarity with how modern deep learning is transforming perception. Key job responsibilities - Design and implement Visual Inertial Odometry algorithms for robust real-time state estimation on robotic platforms like Sprout - Develop multi-sensor fusion pipelines integrating cameras, IMUs, and other sensing modalities for accurate pose tracking - Optimize perception and tracking algorithms for deployment on embedded hardware (e.g., ARM, GPU-accelerated edge devices) under strict latency and power constraints - Apply modern ML-based perception techniques (learned features, depth estimation, neural odometry) to complement and improve classical geometric approaches - Build and maintain calibration, evaluation, and benchmarking infrastructure for perception systems - Collaborate with hardware, controls, and navigation teams to integrate perception outputs into the robot’s autonomy stack - Lead technical projects from research prototyping through production deployment
US, NY, New York
We are seeking a Human-Robot Interaction (HRI) Applied Scientist to develop cutting-edge interactions that make robots feel alive, personal, and fun. In this role, you will focus on verbal and non-verbal conversational systems, social dynamics, memory, and long-term relationship formation between robots, their environments, and the people they interact with. Your contributions will be essential in advancing robotics by enabling expressive, socially intelligent, and trustworthy interactions between robots and humans. Key job responsibilities - Develop interactive systems that leverage large language models, multimodal inputs and outputs, reinforcement learning from human feedback, or other advanced techniques to achieve fluid, engaging, and socially appropriate robot behavior - Design and implement intelligent conversational systems that handle turn-taking, grounding, interruption, and incorporates context drawn from a robot's physical environment and shared history with a user - Integrate perceptual sensor streams including gaze, facial expression, gesture, posture, and more to understand social context and produce coherent, lifelike interactions. - Develop memory and personalization systems that allow robots to form lasting relationships with individual users, learn their environments, and adapt their behavior over weeks and months - Stay updated on advancements in HRI, NLP, multimodal AI, and cognitive and social science to apply cutting-edge techniques to robot interaction challenges - Lead technical projects from conception through production deployment - Mentor junior scientists and engineers - Bridge research initiatives with practical engineering implementation
US, CA, Sunnyvale
Prime Video is a first-stop entertainment destination offering customers a vast collection of premium programming in one app available across thousands of devices. Prime members can customize their viewing experience and find their favorite movies, series, documentaries, and live sports – including Amazon MGM Studios-produced series and movies; licensed fan favorites; and programming from Prime Video add-on subscriptions such as Apple TV+, Max, Crunchyroll and MGM+. All customers, regardless of whether they have a Prime membership or not, can rent or buy titles via the Prime Video Store, and can enjoy even more content for free with ads. Are you interested in shaping the future of entertainment? Prime Video's technology teams are creating best-in-class digital video experience. As a Prime Video technologist, you’ll have end-to-end ownership of the product, user experience, design, and technology required to deliver state-of-the-art experiences for our customers. You’ll get to work on projects that are fast-paced, challenging, and varied. You’ll also be able to experiment with new possibilities, take risks, and collaborate with remarkable people. We’ll look for you to bring your diverse perspectives, ideas, and skill-sets to make Prime Video even better for our customers. With global opportunities for talented technologists, you can decide where a career Prime Video Tech takes you! Key job responsibilities Develop foundation models for content understanding using state-of-the-art deep learning and multimodal learning techniques to analyze video and text Build time sequence foundation models to understand and predict customer behavior patterns and viewing trajectories Work closely with engineers and product managers to design, implement and launch solutions end-to-end across various Prime Video experiences Design and conduct offline and online (A/B) experiments to evaluate proposed solutions based on in-depth data analyses Effectively communicate technical and non-technical ideas with teammates and stakeholders Stay up-to-date with advancements and the latest modeling techniques in foundation models, multimodal learning, and time series analysis Publish your research findings in top conferences and journals A day in the life We're using advanced approaches such as foundation models to connect information about our videos and customers from a variety of information sources, acquiring and processing data sets on a scale that only a few companies in the world can match. This will enable us to recommend titles effectively, even when we don't have a large behavioral signal (to tackle the cold-start title problem). It will also allow us to find our customer's niche interests, helping them discover groups of titles that they didn't even know existed. We are looking for creative & customer obsessed machine learning scientists who can apply the latest research, state of the art algorithms and ML to build highly scalable page personalization solutions. You'll be a research leader in the space and a hands-on ML practitioner, guiding and collaborating with talented teams of engineers and scientists and senior leaders in the Prime Video organization. You will also have the opportunity to publish your research at internal and external conferences. About the team Prime Video Recommendation Science team owns science solution to power recommendation and personalization experience on various Prime Video surfaces and devices. We work closely with the engineering teams to launch our solutions in production.
US, WA, Seattle
Innovators wanted! Are you an entrepreneur? A builder? A dreamer? This role is part of an Amazon Special Projects team that takes the company’s Think Big leadership principle to the nextlevel. We focus on creating entirely new products and services with a goal of positively impacting the lives of our customers. No industries or subject areas are out of bounds. If you’re interested in innovating at scale to address big challenges in the world, this is the team for you. As a Research Scientist, you will work with a unique and gifted team developing exciting products for consumers and collaborate with cross-functional teams. Our team rewards intellectual curiosity while maintaining a laser-focus in bringing products to market. At the intersession of both academic and applied research in this product area, you have the opportunity to work together with some of the most talented scientists, engineers, and product managers. Here at Amazon, we embrace our differences. We are committed to furthering our culture of inclusion. We have thirteen employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We are constantly learning through programs that are local, regional, and global. Amazon’s culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust. Our team highly values work-life balance, mentorship and career growth. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We care about your career growth and strive to assign projects and offer training that will challenge you to become your best.