Model Observatory / GitHub Copilot
Sources & methodology ↗
DEVELOPER INTELLIGENCE

The evidence. Your choice.

Published engineering benchmarks and Copilot prices, in one view.

DATA SNAPSHOT24 Sep 2026Curated snapshot · no automatic refresh

Markers follow your lens and filters. Cheapest: 50k uncached input + 10k output tokens at default rates. Best value: complete lens score ÷ that cost. Excludes utility models, cache charges, seat fees and allowances; not measured cost per completed task.
Published metrics · higher is better
READ THE NUMBERS

Traceable, not interchangeable.

What the scores measure

These are external agent–model evaluations, not tests inside Copilot. A shared benchmark name does not imply identical tools, effort or compute budgets. Details retain published uncertainty and agent labels.

How the company score works

Σ(capability weight × selected benchmark score) ÷ Σ(capability weights). Each capability uses one fixed benchmark, so adding benchmarks does not add votes. Task success rates use percentages; LiveBench Coding uses its 0–100 category score; architecture Node F1 is multiplied by 100. These measure different things. No min–max rescaling. A missing selected result suppresses the score. Versions, subsets and prompt settings stay separate. Review uses the published overall score × 100; false-positive rates are displayed separately and never rewarded. Cost is separate. Equal weights are a starting point, not an approved policy.

What the price means

USD per million tokens at GitHub's default context tier. Input and output are shown separately. Model details include cache and long-context rates. Seat fees, allowances and workload token counts are separate.

GitHub pricing ↗

Catalog covers GitHub’s named supported models plus listed utility models. Actual access depends on plan, client and enterprise policy; your tenant has not been connected. This page does not enable or disable models. Dates marked “not published” are not inferred from crawl dates. Check availability ↗

MODEL EVIDENCE