Pass@k measures the fraction of tasks a model solves in at least one of k attempts. It captures whether a model can succeed given retries, which matters for coding and agent tasks where a single sample may fail by chance. Report pass@1 and pass@k together to separate capability from luck.
Why it matters
Some tasks allow retries - a coding agent can try again. Pass@k measures whether the model can eventually succeed, not just whether it succeeds on the first sample.
How it works
Run the task k times and measure the fraction solved at least once. Pass@1 is single-attempt accuracy; the gap to pass@k shows how much luck versus capability is in the score.