🏆 AI 모델 벤치마크 순위
주요 AI 모델의 순위를 두 관점으로 보여줍니다 — 실사용자 선호(Arena Elo)와 실제 능력(LiveBench 종합점수). 자동 갱신됩니다.
실제 사용자가 두 모델의 답변을 블라인드로 비교 투표해 산출한 Elo 순위입니다. 말투·서식 선호가 반영되며, 능력 정답률과는 다를 수 있습니다.
마지막 업데이트: 2026-09-04 · Arena 발표 기준 Sep 2, 2026
| 순위 | 모델 | 제작사 | 점수(ELO) | 신뢰구간 | 투표수 | 라이선스 |
|---|---|---|---|---|---|---|
| 🥇– | claude-fable-5 | Anthropic | 1,507 | ±5 | 27,189 | 독점 |
| 🥈– | claude-opus-4-6-high | Anthropic | 1,505 | ±4 | 72,099 | 독점 |
| 🥉– | claude-fable-5.1-max | Anthropic | 1,504 | ±11 | 2,906 | 독점 |
| 4▼1 | claude-opus-4-7-high | Anthropic | 1,502 | ±4 | 60,103 | 독점 |
| 5▼1 | muse-spark-1.2 (xHigh) | Meta | 1,499 | ±10 | 3,240 | 독점 |
| 6▼1 | claude-opus-4-6 | Anthropic | 1,498 | ±3 | 76,015 | 독점 |
| 7▼1 | claude-opus-4-7 | Anthropic | 1,494 | ±4 | 61,238 | 독점 |
| 8– | gemini-3.8-flash-high | 1,494 | ±9 | 5,125 | 독점 | |
| 9▼2 | claude-opus-5-high | Anthropic | 1,493 | ±5 | 35,174 | 독점 |
| 10▼2 | muse-spark-1.1 | Meta | 1,492 | ±5 | 24,064 | 독점 |
| 11▼2 | gemini-3.7-flash-high | 1,491 | ±8 | 5,682 | 독점 | |
| 12▼2 | kimi-k3-max | Moonshot | 1,489 | ±5 | 18,090 | |
| 13– | muse-spark | Meta | 1,488 | ±6 | 13,572 | 독점 |
| 14– | claude-opus-5-max | Anthropic | 1,488 | ±6 | 17,119 | 독점 |
| 15– | gemini-3.1-pro-preview | 1,487 | ±3 | 102,999 | 독점 | |
| 16– | gemini-3-pro | 1,486 | ±4 | 40,668 | 독점 | |
| 17– | gpt-5.6-sol-xhigh | OpenAI | 1,483 | ±5 | 23,413 | 독점 |
| 18– | claude-opus-4-8-high | Anthropic | 1,482 | ±4 | 48,727 | 독점 |
| 19– | gpt-5.5-high | OpenAI | 1,482 | ±4 | 63,547 | 독점 |
| 20– | glm-5.3-max | Z.ai | 1,482 | ±7 | 7,668 | 오픈 |
데이터 출처: Arena (구 LMArena) 리더보드. 점수는 사용자 블라인드 비교 투표 기반 Elo이며, 순위는 수시로 변동됩니다.