Pass@k: Measuring Code Generation Success

黎 浩然/ 11 10 月, 2026/ 大语言模型/LARGELANGUAGEMODEL/LLM, 机器学习/MACHINELEARNING, 研究生/POSTGRADUATE, 计算机/COMPUTER/ 0 comments

A code model’s pass@10 can exceed pass@1 simply because it gets nine more attempts. pass@k asks how likely at least one of k candidates for a problem is to pass the tests. It does not establish that the default displayed candidate is correct, or solve the problem of selecting the correct one.

Contents
  1. Count ten candidates for one problem
  2. Recalculate the combinations without running model code
  3. Why not use 1−(1−c/n)ᵏ directly?
  4. Average over problems, not a mixed candidate pool
  5. Read the generation and test conditions
  6. Sources

Count ten candidates for one problem

Suppose ten candidates for an invented problem include two test-passing samples: n=10 and c=2. Candidates may have identical content, but the example counts generated sample indices rather than deduplicating them.

1
Pass
2
Pass
3
Fail
4
Fail
5
Fail
6
Fail
7
Fail
8
Fail
9
Fail
10
Fail
Original candidate-pool illustration: two passing and eight failing samples, with invented labels rather than model outputs. Selecting three indices gives 120 subsets, of which 64 contain a pass.

Selecting three indices uniformly without replacement gives C(10,3)=120 subsets. C(8,3)=56 contain only failing samples. The fraction containing at least one pass is therefore 1−56/120=8/15, approximately 53.33%. This is the pool’s per-problem pass@3 estimate.

The general expression is 1 − C(n−c,k) / C(n,k), requiring n≥k. If fewer than k samples fail, every subset contains a pass and the estimate is one. If c=0, it is zero. At k=1, it reduces to c/n.

Recalculate the combinations without running model code

The original Python example uses exact fractions for this small counting exercise. It was executed to check common k values and the zero-pass boundary. An additional exhaustive enumeration checked all 440 valid (n,c,k) cases for n from one to ten against direct subset counting.

from math import comb
from fractions import Fraction

def pass_estimate(n, c, k):
    if not (n >= 1 and 0 <= c <= n and 1 <= k <= n):
        raise ValueError("require n>=1, 0<=c<=n, 1<=k<=n")
    if n-c < k:
        return Fraction(1,1)
    return 1-Fraction(comb(n-c,k),comb(n,k))

assert pass_estimate(10,2,1) == Fraction(1,5)
assert pass_estimate(10,2,3) == Fraction(8,15)
assert pass_estimate(10,2,5) == Fraction(7,9)
assert pass_estimate(10,0,3) == 0
assert pass_estimate(10,2,10) == 1
for k in [1,3,5,10]:
    print(k, round(float(pass_estimate(10,2,k)),6))
k Per-problem estimate Interpretation
1 20% One randomly selected sample
3 About 53.33% At least one pass among three
5 About 77.78% At least one pass among five
10 100% This observed pool contains passing samples

The last row does not guarantee a correct candidate in the next batch. This code calculates invented labels only: it downloads no model, generates no programs and executes no candidate programs. Larger evaluations need appropriate numerical implementations rather than treating a teaching combination calculator as a production evaluator.

Why not use 1−(1−c/n)ᵏ directly?

If the true single-attempt success probability p is known, k independent attempts succeed at least once with probability 1−(1−p)ᵏ. But c/n is a finite-sample estimate of p, not known p.

Substituting 0.2 and k=3 gives 1−0.8³=0.488 rather than 8/15. One calculation plugs an estimate into a nonlinear expression; the other counts subsets drawn without replacement from the observed pool. They are not interchangeable. The original HumanEval paper analyzes the combination estimator; interpreting it as unbiased relies on independent, identically distributed samples under the same generation settings.

Generating a new answer after reading failure feedback and changing the prompt is a repair workflow. Report it separately rather than hiding a changed process and budget behind the independent-sampling formula.

Average over problems, not a mixed candidate pool

Overall evaluation ordinarily estimates each problem first and averages over problems. Pooling all successes and candidates gives heavily sampled problems more weight when sample counts differ.

For k=1, let problem A have ten samples, all passing, and problem B have one hundred, all failing. The problem average is (1+0)/2=50%; pooling candidates gives 10/110≈9.09%. These numbers answer different questions. Formal comparisons should still keep per-problem generation budgets consistent.

Read the generation and test conditions

The 2021 work “Evaluating Large Language Models Trained on Code” was released as an arXiv preprint alongside HumanEval. This article uses its metric definition, not historical model scores as a current ranking, and reproduces no model experiment.

Record model version, prompt, samples per problem, k, sampling configuration, test version and rules for timeouts and failures. If a product displays one candidate, measure its actual selection strategy separately. Evaluation tests can identify whether a pool contains a pass without showing how users would choose it.

Passing a finite test suite is not proof of correctness for every input. Read pass@k together with task cost, first-attempt usability and failure types to understand what extra generation attempts provide.

中文版

Sources

Evaluating Large Language Models Trained on Code (2021 preprint); HumanEval; Official estimator implementation.

Support

If this article helped you, you can support this site.

WeChat support QR code; click to enlarge
WeChat
Alipay support QR code; click to enlarge
Alipay
Buy Me a Coffee; support this site
Buy Me a Coffee

Click a QR code to enlarge. More options: support page。

Share this Post

Leave a Comment

您的邮箱地址不会被公开。 必填项已用 * 标注

*
*