Customer-obsessed science
Research areas
-
August 26, 20265 min readDiscounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.
-
August 21, 20269 min read
-
July 30, 20268 min read
-
-
July 9, 202610 min read
Featured news
-
The 11th Workshop on Financial Technology and Natural Language Processing2026Long-context financial document understanding and reasoning pose significant challenges for small language models (SLMs). In this paper, we scale the financial long-context reasoning capability of SLMs through reinforcement learning. Specifically, we propose an efficient curriculum reinforcement learning recipe that features staged training across context lengths and difficulty-aware data sampling. To support
-
2026Optical Character Recognition (OCR) is a fundamental task for digitizing information, serving as a critical bridge between visual data and textual understanding. While modern Vision-Language Models (VLM) have achieved high accuracy in this domain, they predominantly rely on autoregressive decoding, which becomes computationally expensive and slow for long documents as it requires a sequential forward pass
-
2026Creating authentic agentic twins of mobile users through realistic user behavior simulation is critical to truly understand and anticipate customer needs at scale. Toward this, we introduce Agentic Twins of Mobile Users (AgenTwin), an end-to-end framework for high-fidelity mobile user simulation that addresses four key limitations in existing approaches: subjective decision-making diversity, scalable experience
-
EuroSys 20262026Modern large language models rely on attention mechanisms that attend to all tokens in a sequence, resulting in quadratic computational complexity that limits scalability. While sparse attention reduces compute and memory requirements by attending to only important tokens, implementing these techniques presents significant challenges due to the complexity of combining static and dynamic sparse patterns
-
CIKM 20262026Generative, LLM-powered search demos beautifully, then collides with production: hard latency deadlines, a bill that scales with traffic, and near-zero tolerance for failure. This experience report argues that the hidden specification for production AI search is the customer's expectation of difficulty—how hard a request looks to them—which sets both the quality they demand and the latency they tolerate
Collaborations
View allWhether you're a faculty member or student, there are number of ways you can engage with Amazon.
View all