LILT Launches AURORA, the First Multilingual AI Leaderboard That Measures Frontier Models on Non-English Enterprise Agentic Tasks Grounded in Language and Culture

LILT Launches AURORA, the First Multilingual AI Leaderboard That Measures Frontier Models on Non-English Enterprise Agentic Tasks Grounded in Language and Culture

PR Newswire

SAN FRANCISCO, Sept. 30, 2026 /PRNewswire/ — LILT today launched AURORA, the industry’s first Multilingual AI Leaderboard that measures frontier models on non-English agentic, multimodal, and socio-cultural tasks through scientifically rigorous evaluation.

Lilt is an AI-powered enterprise language translation company on a mission to make the world’s information accessible to everyone regardless of where they were born or which language they speak. Visit us online at www.lilt.com.

Enterprises are deploying agents worldwide, yet nearly every published measure of frontier progress relies on English-centric or translated benchmarks. With agents doing customer-facing work, performance in every language now carries commercial weight.

“The industry is choosing models on a scoreboard that stops at English. That gap used to cost you an awkward translation. Now that agents are taking real actions in the real world, it costs you a wrong decision, and you find out from your users.” — Spence Green, CEO and co-founder, LILT

Measuring Multilingual Performance:

AURORA addresses the multilingual gap by testing models against LILT’s multilingual benchmark suite featuring tasks designed and verified by native-language domain experts. These tasks represent real-world enterprise applications, such as software development, customer support, and complex workflows, each grounded in language, region and culture.

Specifically, AURORA provides visibility across LILT’s multilingual benchmarks, including:

  • Multilingual Terminal-bench: Coding tasks representing challenges in software written for native-language users.
  • Multilingual τ³-bench: Multi-turn customer support in airline, telecom, retail, and banking.
  • Multilingual MultiChallenge: Long-context instruction-following, memory and self-coherence.
  • Multilingual GAIA-v2-LILT: Agentic reasoning and tool use.

AURORA is developed and managed by LILT’s Applied AI practice, a PhD-led research team with 10+ years’ experience in multilingual AI.

Their latest analysis shows model quality can differ substantially between languages. For example, in coding, GPT 5.5 performs best in Spanish, Claude Opus 5.5 wins in Japanese, while Muse Spark 1.3 leads in Serbian.

Availability

AURORA is live at https://aurora.lilt.com/, where readers can compare models by language and task. AI teams interested in private, custom benchmarks can contact LILT at contact@lilt.com.

About LILT

LILT is the leading agentic AI solution to make anything multilingual at scale for enterprise, public sector, and Frontier labs.  LILT’s Applied AI services give AI teams private multilingual benchmarks, evaluation, and training data needed to ship better agents and models. Leading Frontier labs and organizations like NVIDIA, Intel, and L’Oréal rely on LILT to expand global reach. Learn more at lilt.com.

Cision View original content to download multimedia:https://www.prnewswire.com/news-releases/lilt-launches-aurora-the-first-multilingual-ai-leaderboard-that-measures-frontier-models-on-non-english-enterprise-agentic-tasks-grounded-in-language-and-culture-302894528.html

SOURCE LILT