Skip to content
ai-supply.store
खोजेंश्रेणियाँलीडरबोर्डसमुदायAgent APIFAQ
प्रकाशित करेंसाइन इन
catalog / Language & NLP / OpenOrca
▣DatasetLanguage & NLPFree

OpenOrca

MIT-licensed 4.2M instruction dataset — GPT-4/3.5 augmented CoT traces that power top open-source fine-tunes.

@ai-supply
इंस्टॉल98k
रेटिंग★ 4.6
समीक्षाएं33
↗ सोर्स रिपॉज़िटरी

OpenOrca

OpenOrca is a 4.2-million-row instruction fine-tuning dataset released under the MIT license by the Open-Orca community. It replicates and extends the Microsoft Orca research paper by generating chain-of-thought (CoT) explanation traces using GPT-4 and GPT-3.5, then attaching them to the FLAN collection. Models fine-tuned on OpenOrca consistently outperform same-size alternatives on reasoning and instruction-following tasks.

Key features

  • 4.2M rows of system prompt + instruction + GPT-4/3.5 CoT response triples
  • Covers FLAN-1M, CoT, and Open Platypus subsets
  • MIT license — fully permissive for commercial fine-tuning
  • Powers open-weight models like OpenHermes, MistralOrca, LLaMA-Orca
  • Parquet format — fast loading with HuggingFace datasets

Quick start

from datasets import load_dataset

# Load a 10% sample to start
ds = load_dataset("Open-Orca/OpenOrca", split="train[:10%]")
print(ds[0])
# {'system_prompt': ..., 'question': ..., 'response': ...}

# Filter GPT-4 only rows for highest quality
gpt4_only = ds.filter(lambda x: "gpt-4" in x["id"])
print(f"{len(gpt4_only)} GPT-4 rows")

Install via ai-supply

npx ai-supply add open-orca-dataset

Curated mirror of the open-source OpenOrca (MIT). Get it from the source.

More from @ai-supply

View profile →
◐Model
llama.cpp
Pure C/C++ LLM inference library — run quantized models on CPU, Metal, CUDA and more.
↓ 900k★ 4.9
⇄Connector
vLLM
High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.
↓ 820k★ 4.9
◉Agent
MetaGPT
Multi-agent framework that assigns GPT roles (PM, engineer, QA) to solve complex software tasks end-to-end.
↓ 820k★ 4.8
◆Skill
NLTK
The Natural Language Toolkit — Python's foundational NLP library for tokenization, POS tagging, parsing, and corpora.
↓ 760k★ 4.7
ai-supply.store

AI क्षमताओं का मार्केटप्लेस। स्किल्स, MCP सर्वर, प्लगइन्स, एजेंट, डेटासेट — मानवों द्वारा खोजने योग्य, मशीनों द्वारा उपभोग योग्य।

api · v3.1status · all green
संपर्क करें
support@ai-supply.storesecurity@ai-supply.store
मार्केटप्लेस
  • खोजें
  • श्रेणियाँ
  • लीडरबोर्ड
  • बेंचमार्क
समुदाय
  • समुदाय
  • FAQ
एजेंट के लिए
  • क्विकस्टार्ट (60s)
  • एजेंट अधिकृत करें
  • Agent API
  • OpenAPI स्पेसिफिकेशन
बिल्डर्स के लिए
  • प्रकाशित करें
  • डैशबोर्ड
  • राजस्व हिस्सेदारी
खाता
  • साइन इन
  • सेटिंग्स
कानूनी
  • नियम व शर्तें
  • प्रकाशक अनुबंध
  • स्वीकार्य उपयोग नीति
  • गोपनीयता