Skip to content
ai-supply.store
ExplorarCategoríasClasificacionesComunidadAgent APIFAQ
PublicarIniciar sesión
catalog / Language & NLP / OpenOrca
▣DatasetLanguage & NLPFree

OpenOrca

MIT-licensed 4.2M instruction dataset — GPT-4/3.5 augmented CoT traces that power top open-source fine-tunes.

@ai-supply
Instalaciones98k
Valoración★ 4.6
Reseñas33
↗ Repositorio fuente

OpenOrca

OpenOrca is a 4.2-million-row instruction fine-tuning dataset released under the MIT license by the Open-Orca community. It replicates and extends the Microsoft Orca research paper by generating chain-of-thought (CoT) explanation traces using GPT-4 and GPT-3.5, then attaching them to the FLAN collection. Models fine-tuned on OpenOrca consistently outperform same-size alternatives on reasoning and instruction-following tasks.

Key features

  • 4.2M rows of system prompt + instruction + GPT-4/3.5 CoT response triples
  • Covers FLAN-1M, CoT, and Open Platypus subsets
  • MIT license — fully permissive for commercial fine-tuning
  • Powers open-weight models like OpenHermes, MistralOrca, LLaMA-Orca
  • Parquet format — fast loading with HuggingFace datasets

Quick start

from datasets import load_dataset

# Load a 10% sample to start
ds = load_dataset("Open-Orca/OpenOrca", split="train[:10%]")
print(ds[0])
# {'system_prompt': ..., 'question': ..., 'response': ...}

# Filter GPT-4 only rows for highest quality
gpt4_only = ds.filter(lambda x: "gpt-4" in x["id"])
print(f"{len(gpt4_only)} GPT-4 rows")

Install via ai-supply

npx ai-supply add open-orca-dataset

Curated mirror of the open-source OpenOrca (MIT). Get it from the source.

More from @ai-supply

View profile →
◐Model
llama.cpp
Pure C/C++ LLM inference library — run quantized models on CPU, Metal, CUDA and more.
↓ 900k★ 4.9
⇄Connector
vLLM
High-throughput, memory-efficient LLM inference engine with PagedAttention and continuous batching.
↓ 820k★ 4.9
◉Agent
MetaGPT
Multi-agent framework that assigns GPT roles (PM, engineer, QA) to solve complex software tasks end-to-end.
↓ 820k★ 4.8
◆Skill
NLTK
The Natural Language Toolkit — Python's foundational NLP library for tokenization, POS tagging, parsing, and corpora.
↓ 760k★ 4.7
ai-supply.store

El marketplace de capacidades de IA. Habilidades, MCPs, plugins, agentes, datasets — descubribles por humanos, consumibles por máquinas.

api · v3.1status · all green
Contacto
support@ai-supply.storesecurity@ai-supply.store
Marketplace
  • Explorar
  • Categorías
  • Clasificaciones
  • Benchmarks
Comunidad
  • Comunidad
  • FAQ
Para agentes
  • Inicio rápido (60s)
  • Autorizar un agente
  • Agent API
  • Especificación OpenAPI
Para desarrolladores
  • Publicar
  • Panel
  • Reparto de ingresos
Cuenta
  • Iniciar sesión
  • Configuración
Legal
  • Términos
  • Acuerdo de editor
  • Uso aceptable
  • Privacidad