Pass@k: Measuring Code Generation Success
A code model’s pass@10 can exceed pass@1 simply because it gets nine more attempts. pass@k asks how likely at least one of k candidates for a problem is to pass the tests. It does not establish that the default displayed candidate is correct, or solve the problem of selecting the correct one.
Contents
Count ten candidates for one problem
Suppose ten candidates for an invented problem include two test-passing samples: n=10 and c=2. Candidates may have identical content, but the example counts generated sample indices rather than deduplicating them.
Pass
Pass
Fail
Fail
Fail
Fail
Fail
Fail
Fail
Fail
Selecting three indices uniformly without replacement gives C(10,3)=120 subsets. C(8,3)=56 contain only failing samples. The fraction containing at least one pass is therefore 1−56/120=8/15, approximately 53.33%. This is the pool’s per-problem pass@3 estimate.
The general expression is 1 − C(n−c,k) / C(n,k), requiring n≥k. If fewer than k samples fail, every subset contains a pass and the estimate is one. If c=0, it is zero. At k=1, it reduces to c/n.
Recalculate the combinations without running model code
The original Python example uses exact fractions for this small counting exercise. It was executed to check common k values and the zero-pass boundary. An additional exhaustive enumeration checked all 440 valid (n,c,k) cases for n from one to ten against direct subset counting.
from math import comb
from fractions import Fraction
def pass_estimate(n, c, k):
if not (n >= 1 and 0 <= c <= n and 1 <= k <= n):
raise ValueError("require n>=1, 0<=c<=n, 1<=k<=n")
if n-c < k:
return Fraction(1,1)
return 1-Fraction(comb(n-c,k),comb(n,k))
assert pass_estimate(10,2,1) == Fraction(1,5)
assert pass_estimate(10,2,3) == Fraction(8,15)
assert pass_estimate(10,2,5) == Fraction(7,9)
assert pass_estimate(10,0,3) == 0
assert pass_estimate(10,2,10) == 1
for k in [1,3,5,10]:
print(k, round(float(pass_estimate(10,2,k)),6))
| k | Per-problem estimate | Interpretation |
|---|---|---|
| 1 | 20% | One randomly selected sample |
| 3 | About 53.33% | At least one pass among three |
| 5 | About 77.78% | At least one pass among five |
| 10 | 100% | This observed pool contains passing samples |
The last row does not guarantee a correct candidate in the next batch. This code calculates invented labels only: it downloads no model, generates no programs and executes no candidate programs. Larger evaluations need appropriate numerical implementations rather than treating a teaching combination calculator as a production evaluator.
Why not use 1−(1−c/n)ᵏ directly?
If the true single-attempt success probability p is known, k independent attempts succeed at least once with probability 1−(1−p)ᵏ. But c/n is a finite-sample estimate of p, not known p.
Substituting 0.2 and k=3 gives 1−0.8³=0.488 rather than 8/15. One calculation plugs an estimate into a nonlinear expression; the other counts subsets drawn without replacement from the observed pool. They are not interchangeable. The original HumanEval paper analyzes the combination estimator; interpreting it as unbiased relies on independent, identically distributed samples under the same generation settings.
Generating a new answer after reading failure feedback and changing the prompt is a repair workflow. Report it separately rather than hiding a changed process and budget behind the independent-sampling formula.
Average over problems, not a mixed candidate pool
Overall evaluation ordinarily estimates each problem first and averages over problems. Pooling all successes and candidates gives heavily sampled problems more weight when sample counts differ.
For k=1, let problem A have ten samples, all passing, and problem B have one hundred, all failing. The problem average is (1+0)/2=50%; pooling candidates gives 10/110≈9.09%. These numbers answer different questions. Formal comparisons should still keep per-problem generation budgets consistent.
Read the generation and test conditions
The 2021 work “Evaluating Large Language Models Trained on Code” was released as an arXiv preprint alongside HumanEval. This article uses its metric definition, not historical model scores as a current ranking, and reproduces no model experiment.
Record model version, prompt, samples per problem, k, sampling configuration, test version and rules for timeouts and failures. If a product displays one candidate, measure its actual selection strategy separately. Evaluation tests can identify whether a pool contains a pass without showing how users would choose it.
Passing a finite test suite is not proof of correctness for every input. Read pass@k together with task cost, first-attempt usability and failure types to understand what extra generation attempts provide.
Sources
Evaluating Large Language Models Trained on Code (2021 preprint); HumanEval; Official estimator implementation.
Support
If this article helped you, you can support this site.
Click a QR code to enlarge. More options: support page。


