Universal LLM Rankings

See individual leaderboards
# RecommendationsSegmentationE-commerce UnderstandingBrazilian KnowledgeTaxonomy Prediction
1
Made by ProsusLCM-3 Picanha27B
Prosus $0.90 est. 70.9 67.7 66.4 64.1 84.4 72.1
2
Qwen3.8-2.4T-A95B2400B
Alibaba $6.00 69.7 62.8 58.8 62.1 88.0 76.7
3
GPT-5.6 Terra
OpenAI $12.00 69.4 61.8 64.8 61.3 86.6 72.4
4
GPT-5.6 Sol
OpenAI $20.00 68.9 64.8 60.0 64.4 88.5 66.9
5
GPT-5.6 Luna
OpenAI $1.20 68.6 60.9 66.6 59.4 85.9 70.4
6
Made by ProsusLCM-3 Feijoada9B
Prosus $0.29 est. 68.3 67.8 67.6 57.1 79.3 69.9
7
Claude Sonnet 5
Anthropic $10.00 68.2 57.7 64.6 60.7 86.7 71.2
8
Claude Opus 5
Anthropic $25.00 67.7 63.8 61.8 60.7 88.8 63.4
9
DeepSeek-V4.1-Flash552B
DeepSeek $1.20 66.4 59.4 62.0 56.9 84.5 69.1
10
GLM-5.3-Flash320B
Z.ai $0.50 66.4 57.6 61.4 60.1 84.4 68.3
11
GLM-5.3753B
Z.ai $4.40 65.0 59.3 47.8 60.1 86.2 71.3
12
Qwen3.8 27B27B
Alibaba $3.00 62.9 60.0 60.2 61.5 80.5 52.0
13
Qwen3.5 9B9B
Alibaba $0.25 54.8 50.6 52.4 56.2 74.7 39.9
Rows

13 models

Green = Top result Prosus LCM

How scores work. Every score is rescaled to 0–100 using each metric's original range. For metrics where lower is better the scale is flipped so 100 is always best. The overall score is calculated as the unweighted average of all columns. Missing scores are left out of the average rather than counted as zero.

Because of the rescaling the scores displayed differ from the raw results shown on individual leaderboards. To see a detailed breakdown of benchmark scores in their original format, click "Expand all".

How prices work. The $/1M tokens column is the price of a million output tokens: the published output rate for API models, and for self-hosted models an estimate from the hourly cost of one GPU, its peak throughput and an average throughput utilization.