Customer-obsessed science
Research areas
-
October 1, 202610 min readAugmenting a network graph with agentic AI produces a “digital twin” that can help isolate network failures.
-
-
August 21, 20269 min read
-
July 30, 20268 min read
-
July 29, 20266 min read
Featured news
-
2026Matching users to interest categories at scale is central to personalized shopping, but the task is challenging in large e-commerce platforms, where label spaces continually evolve and user-interest signals are sparse and long tailed. Autoregressive language models are appealing because their world knowledge and semantic priors over descriptors generalize across extreme label spaces and accommodate multiple
-
2026Memory makes LLM-based web agents personalized, powerful, yet exploitable. By storing past interactions to personalize future tasks, agents inadvertently create a persistent attack surface that spans websites and sessions. While existing security research on memory assumes attackers can directly inject into memory storage or exploit shared memory across users, we present a more realistic threat model: contamination
-
2026LLM judges are increasingly used to evaluate open-ended responses, but their scores depend strongly on the rubrics that condition them. A vague rubric asking for a response to be “helpful and factual” can reward polished answers that invent facts or violate user intent. We treat reusable rubrics as measurement specifications: changing the rubric changes the response quality measurement induced by a fixed
-
2026Music recommendations rely on robust artist embeddings that capture stylistic and behavioral similarity. A common approach is to learn track-level embeddings from customer co-listening data and aggregate them per artist, but simple heuristic aggregation (e.g., averaging) treats all tracks equally and may fail to capture which tracks best define an artist’s identity. We introduce MusicContainerNet (MCN),
-
2026Music recommendation surfaces need human-readable explanations, but LLM-quality generation does not scale on the serving path. Using MyMix, an algorithmically-generated personalized playlist surface, as a testbed, we start from a whole-playlist 4B baseline and identify two limitations: feasibility (per-request inference does not scale) and input granularity (the model sees only coarse aggregated tags, not
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all