Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
July 29, 20266 min read
-
Featured news
-
2026Large language models increasingly need to generate structured outputs that conform to predefined schemas, with one common constraint being selection from a finite set of valid strings. Current constrained decoding systems handle this through general-purpose grammar compilation, which becomes prohibitively slow as the number of valid values grows into the thousands, a cardinality wall. We introduce the
-
2026Domain-specific finetuning is essential for dense retrievers, yet not all data pairs contribute equally to the learning process. We introduce OPERA1 , a data pruning framework that exploits this heterogeneity to improve both the effectiveness and efficiency of retrieval model adaptation. We first investigate static pruning (SP), which retains only high-similarity query document pairs, revealing an intrinsic
-
EMNLP 20262026Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, producing prompts up to 3×longer yet no more accurate. We trace this to three deficiencies - incomplete error observation, limited search diversity, and unreliable selection - and propose ESPO (Error-Structured Prompt Optimization), which decomposes prompt optimization into three phases: Diagnose
-
EMNLP 20262026Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates to millions of annotations, rendering human labeling prohibitively costly. While recent work has demonstrated synthetic label generation using LLMs (Negri et al., 2025), deploying such
-
2026Agent evaluation today depends on per-trace LLM-judge inference or human review, too expensive to run on every trace; production systems fall back to sampling a fraction. We find that failing agents leave a detectable behavioral signature in standard observability telemetry: disproportionate effort relative to outcome. We formalize four behavioral failure signatures from agent telemetry, validated on 4,671
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all