What the scores measure
These are external agent–model evaluations, not tests inside Copilot. A shared benchmark name does not imply identical tools, effort or compute budgets. Details retain published uncertainty and agent labels.
Published engineering benchmarks and Copilot prices, in one view.
These are external agent–model evaluations, not tests inside Copilot. A shared benchmark name does not imply identical tools, effort or compute budgets. Details retain published uncertainty and agent labels.
Σ(capability weight × selected benchmark score) ÷ Σ(capability weights). Each capability uses one fixed benchmark, so adding benchmarks does not add votes. Task success rates use percentages; LiveBench Coding uses its 0–100 category score; architecture Node F1 is multiplied by 100. These measure different things. No min–max rescaling. A missing selected result suppresses the score. Versions, subsets and prompt settings stay separate. Review uses the published overall score × 100; false-positive rates are displayed separately and never rewarded. Cost is separate. Equal weights are a starting point, not an approved policy.
USD per million tokens at GitHub's default context tier. Input and output are shown separately. Model details include cache and long-context rates. Seat fees, allowances and workload token counts are separate.
GitHub pricing ↗Catalog covers GitHub’s named supported models plus listed utility models. Actual access depends on plan, client and enterprise policy; your tenant has not been connected. This page does not enable or disable models. Dates marked “not published” are not inferred from crawl dates. Check availability ↗