Skip to content
ai-supply.store
DiscoverCategoriesLeaderboardsCommunityAgent APIFAQ
Sign inSign up free
catalog / Language & NLP / OpenOrca
▣DatasetLanguage & NLPFree

OpenOrca

MIT-licensed 4.2M instruction dataset — GPT-4/3.5 augmented CoT traces that power top open-source fine-tunes.

@ai-supply
Installs116k
↗ Source repository
← More Language & NLPLanguage & NLP leaderboard →How we grade security →Source ↗

OpenOrca

OpenOrca is a 4.2-million-row instruction fine-tuning dataset released under the MIT license by the Open-Orca community. It replicates and extends the Microsoft Orca research paper by generating chain-of-thought (CoT) explanation traces using GPT-4 and GPT-3.5, then attaching them to the FLAN collection. Models fine-tuned on OpenOrca consistently outperform same-size alternatives on reasoning and instruction-following tasks.

Key features

  • 4.2M rows of system prompt + instruction + GPT-4/3.5 CoT response triples
  • Covers FLAN-1M, CoT, and Open Platypus subsets
  • MIT license — fully permissive for commercial fine-tuning
  • Powers open-weight models like OpenHermes, MistralOrca, LLaMA-Orca
  • Parquet format — fast loading with HuggingFace datasets

Quick start

from datasets import load_dataset

# Load a 10% sample to start
ds = load_dataset("Open-Orca/OpenOrca", split="train[:10%]")
print(ds[0])
# {'system_prompt': ..., 'question': ..., 'response': ...}

# Filter GPT-4 only rows for highest quality
gpt4_only = ds.filter(lambda x: "gpt-4" in x["id"])
print(f"{len(gpt4_only)} GPT-4 rows")

Install via ai-supply

npx ai-supply add open-orca-dataset

Curated mirror of the open-source OpenOrca (MIT). Get it from the source.

Rating rank
#1
of 30 in Language & NLP
Install rank
#10
of 30 in Language & NLP
Installs
116k
cat avg 145k
This listing vs category average
Installs
this
cat avg
Adoption trend
See the Language & NLP leaderboard →
Not yet scanned — this listing is queued for source analysis.

Consume OpenOrca programmatically. Authenticate with an API key or session — see Authorize an agent.

# Agents: CHECK BEFORE YOU INSTALL (no auth) — score, grade, level, capability manifest
curl https://ai-supply.store/api/v1/trust/open-orca-dataset

# Gate against your org policy (returns { pass, violations })
curl -X POST https://ai-supply.store/api/v1/trust/open-orca-dataset/check \
  -H "Content-Type: application/json" \
  -d '{"minGrade":"B","denyPermissions":["shell"],"denyUnknownEgress":true}'

# CLI
npx ai-supply add open-orca-dataset

# REST (install → download)
curl -X POST https://ai-supply.store/api/v1/listings/open-orca-dataset/install \
  -H "Authorization: Bearer $AIM_KEY"

# MCP tool
install_listing({ "slug": "open-orca-dataset" })
OpenAPI spec →
vlatest
1mo ago

Curated mirror — latest upstream source. See the repository for tagged releases.

Sign in and install this listing to leave a review.

More from @ai-supply

View profile →
◉Agent
MetaGPT
Multi-agent framework that assigns GPT roles (PM, engineer, QA) to solve complex software tasks end-to-end.
↓ 1.0M
⇄Connector
vLLM
High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.
↓ 892k
⇄Connector
Meilisearch
Lightning-fast open-source search engine with typo-tolerance, semantic hybrid search, and sub-50ms response times.
↓ 811k
△Eval
Weights & Biases (wandb)
ML experiment tracking and visualization — log metrics, hyperparameters, models, and media in real time.
↓ 784k
ai-supply.store

Free, security-vetted AI capabilities — skills, MCPs, plugins, agents, datasets and more, each graded and freshness-tracked, and built for humans and agents alike.

api · v3.1status · all green
Contact
support@ai-supply.storesecurity@ai-supply.store
Catalog
  • Discover
  • Categories
  • Leaderboards
  • Benchmarks
  • Security
  • Scan a repo
Community
  • Community
  • FAQ
For agents
  • Quickstart (60s)
  • Authorize an agent
  • Agent API
  • OpenAPI spec
For builders
  • Publish
  • Dashboard
Account
  • Create account
  • Sign in
  • Settings
Legal
  • Terms
  • Publisher Agreement
  • Acceptable Use
  • Privacy