catalog / Coding / HumanEval
DatasetCodingFree

HumanEval

OpenAI's 164-problem Python benchmark for evaluating code-generation correctness (pass@k).

安装量3.6k
⟳ upstream master@6d43fb9 · updated 1y ago
源代码仓库
! Grade B · 88/100 · ReviewSecurity assessment
No compromise signals2capabilities surfaced1known CVE9of 20 OWASP controls clear
Suspicious code patternsSuspicious code patternsVulnerable dependencies
scanned 1mo agoosv · gitleaks · opengrep · picklescan + heuristicsfull breakdown in the Security tab ↓

HumanEval

OpenAI's HumanEval benchmark — 164 hand-written Python programming problems, each with a function signature, docstring, and unit tests — used to evaluate the functional correctness of code-generation models (the pass@k metric).

The repository includes the problem dataset and an execution harness that runs generated solutions against the hidden tests.

MIT licensed and tiny; a standard reference dataset for measuring code-synthesis capability.

More from @ai-supply

View profile →
Agent
MetaGPT
Multi-agent framework that assigns GPT roles (PM, engineer, QA) to solve complex software tasks end-to-end.
1.0M
Connector
vLLM
High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.
892k
Connector
Meilisearch
Lightning-fast open-source search engine with typo-tolerance, semantic hybrid search, and sub-50ms response times.
811k
Eval
Weights & Biases (wandb)
ML experiment tracking and visualization — log metrics, hyperparameters, models, and media in real time.
784k