AI Leaderboard Arena Raises $200M Series B at $3.1B, Nearly Doubling in 10 Months, and Previews an Alignment Index

On October 8 Arena (formerly LMArena), the crowdsourced AI model evaluation platform, announced a $200 million Series B at a $3.1 billion valuation, led by Lightspeed Venture Partners and Khosla Ventures, with new investors including Salesforce Ventures and existing backers such as a16z and Felicis; its January Series A was at a $1.7 billion post-money valuation. Arena says it has passed $100 million in annualized revenue, with 350 million sessions and 62 million votes on the platform. It also released a preview of the Arena Alignment Index, which starts with three signals verifiable in real agent traces: unauthorized action, false attribution and deceptive completion, covering more than 20 frontier models. According to TechCrunch, OpenAI models top the preliminary ranking, with Claude Opus 5.5 in sixth.

Funding and business

On October 8 Arena (formerly LMArena), the crowdsourced AI model evaluation platform, announced a $200 million Series B at a $3.1 billion valuation. Lightspeed Venture Partners and Khosla Ventures led, joined by new investors Salesforce Ventures, 01 Advisors, Dell Technologies Capital and Endeavor Catalyst, with existing backers including a16z and Felicis. It closed a $150 million Series A at a $1.7 billion post-money valuation only in January, so its valuation nearly doubled in under 10 months.

Arena began in 2023 as a UC Berkeley research project: users enter a prompt, two anonymous models answer, and users vote for the better one. Beyond that free platform, it launched a paid evaluation product for model labs and enterprises last September. Arena says it has passed $100 million in annualized revenue, with 350 million sessions and 62 million votes on the platform, and 7 million sessions in Agent Arena in under five months since launch, all company figures.

From "which is stronger" to "can you trust it"

The same day Arena released a preview of the Arena Alignment Index, starting with three signals verifiable in real agent traces: unauthorized action (doing things the user didn't ask for), false attribution (attributing to the user a claim that their evidence contradicts) and deceptive completion (saying a task is done when it isn't). The definitions draw on those used in OpenAI and Anthropic system cards, and the first results cover more than 20 frontier models. According to TechCrunch, OpenAI models hold the top spots on the preliminary ranking, with Claude Opus 5.5 in sixth.

What it means

Arena rankings often appear in model launch marketing, and its rising valuation shows that "neutral evaluation" has itself become big business. But it also sells evaluation services to the labs whose models it ranks, so questions about its independence will persist. The alignment index targets a real pain point of the agent era: users care whether an agent quietly overstepped or claimed to finish work it didn't. It is still a preview with only three signals, though, and rankings may shift as the methodology changes, so treat it as a reference for model selection, not a verdict.

via: Arena announcement, TechCrunch report