Oumi
An open, end-to-end platform for fine-tuning, evaluating, and deploying foundation models from 10M to 400B+ parameters.
- Source repo
- oumi-ai/oumi
- Stars
- ★ 9.4k
- Last updated
- today
- License
- Apache-2.0
- Primary language
- Python
- FA score
- 58/100 · Major gaps
At a glance
- How it runs
- Works with
- Universal · cross-platformClaude Code · OpenAI APIChatGPT · Claude API (Partial support)
- Cost
- Free software; you pay for model usage
- Setup effort
- High · needs real infrastructure
- You'll need
- Typical use
- A team that wants to fine-tune Qwen3 or a Gemma model for its own domain and would rather start from an existing YAML recipe than write a training script.
- Not a fit if
- Users who only want a chat UI and no config files or CLI
- Teams with no GPU access doing only light experimentation
- Production buyers who need a commercial SLA and vendor support
- Source review
- 58/100 · Major gaps
What does this agent do, and when should you use it?
Oumi is a fully open-source platform that covers the full lifecycle of foundation models: data synthesis and curation, pre-training and fine-tuning, evaluation, inference, and deployment. It is driven by a single oumi CLI plus YAML recipes, so you rarely have to write your own training loop or data pipeline. Training supports SFT, LoRA, QLoRA, DPO, and GRPO, scaling through FSDP, DeepSpeed, and DDP; inference integrates vLLM and SGLang, and models can be pushed to dedicated endpoints with oumi deploy. Beyond text models it handles vision-language models and agentic, tool-using models: you can define executable tool environments (database, HTTP endpoint, lookup, simulated) and run verl GRPO reinforcement learning over them. Evaluation spans standard benchmarks plus multi-criteria RubricJudge and LLM-as-a-Judge. It runs from a laptop up to clusters and clouds via oumi launch (AWS, Azure, GCP, Lambda, Modal, Slurm), and oumi-mcp exposes an MCP server for Claude and Cursor.
Oumi centers on the oumi command-line interface. You author or reuse a YAML config from configs/recipes and run oumi train -c <config> to launch training (SFT, LoRA, QLoRA, DPO, GRPO on PyTorch, optionally with FSDP/DeepSpeed/DDP), oumi evaluate -c <config> to score a model on standard benchmarks or custom judges, and oumi infer -c <config> --interactive for interactive inference through engines such as vLLM, SGLang, Fireworks, or OpenRouter. On the data side it reads raw corpora and uses LLM judges, including the multi-criteria RubricJudge, to synthesize and filter training data, including multi-turn tool-use conversations. For agentic training it manages executable tool environments (database, HTTP endpoint, lookup, simulated) and applies verl GRPO over them. For delivery, oumi deploy publishes models to dedicated inference endpoints on Fireworks and Parasail, oumi launch up submits jobs to cloud platforms, and oumi-mcp runs an MCP server consumable by Claude and Cursor. oumi analyze covers analysis, and the configs/recipes tree ships ready-made configurations for Qwen, Gemma, Llama, DeepSeek, GLM, Phi, OLMo, gpt-oss, Falcon, and more.
- A team that wants to fine-tune Qwen3 or a Gemma model for its own domain and would rather start from an existing YAML recipe than write a training script.
- A research group running reproducible GRPO reinforcement-learning experiments across multiple GPUs or a cloud cluster, needing FSDP or DeepSpeed.
- A developer training a tool-calling agent model who needs to construct database, HTTP, or lookup tool environments and do reinforcement learning over them.
- A data team that must synthesize, filter, and curate training data at scale with LLM-as-a-Judge or RubricJudge, including vision-language datasets.
- An engineering team deploying a fine-tuned model locally through vLLM/SGLang, or publishing it to a managed endpoint with oumi deploy.
- Developers already working inside an AI-assisted IDE such as Claude or Cursor who want to drive Oumi workflows through the oumi-mcp MCP server.
How do you install or deploy this agent?
The README recommends pip (uv):
# Basic installation
uv pip install oumi
# With GPU support
uv pip install 'oumi[gpu]'
# Latest development version
uv pip install git+https://github.com/oumi-ai/oumi.gitDocker and an experimental one-line script are also documented:
docker pull ghcr.io/oumi-ai/oumi:latest
docker run --gpus all -it ghcr.io/oumi-ai/oumi:latest oumi --helpcurl -LsSf https://oumi.ai/install.sh | bashOumi is in beta and under active development: core features are stable, but some advanced features may change. More installation options are in the official installation guide.
How do you use this agent?
Pick a YAML config under configs/recipes and drive it with the oumi CLI:
# Training
oumi train -c configs/recipes/smollm/sft/135m/quickstart_train.yaml
# Evaluation
oumi evaluate -c configs/recipes/smollm/evaluation/135m/quickstart_eval.yaml
# Inference
oumi infer -c configs/recipes/smollm/inference/135m_infer.yaml --interactiveTo run remotely, submit with oumi launch; to train inside the container, mount your working directory and pass --config:
oumi launch up -c configs/recipes/smollm/sft/135m/quickstart_gcp_job.yaml --resources.cloud awsdocker run --gpus all -v $(pwd):/workspace -it ghcr.io/oumi-ai/oumi:latest \
oumi train --config /workspace/my_config.yamlWhat are this agent's strengths and limitations?
- One CLI plus YAML recipes covers the whole pipeline — training, evaluation, inference, data synthesis, and deployment — with no custom training loop.
- Broad training method coverage: SFT, LoRA, QLoRA, DPO, and GRPO, with native FSDP, DeepSpeed, and DDP support for scaling.
- Agentic focus backed by concrete components: executable tool environments (database, HTTP endpoint, lookup, simulated) and verl GRPO reinforcement learning over them.
- Evaluation is first-class, combining standard benchmarks with multi-criteria RubricJudge and LLM-as-a-Judge.
- Multiple delivery paths: vLLM/SGLang inference, oumi deploy to dedicated endpoints, oumi-mcp for Claude and Cursor, and oumi launch to AWS, Azure, GCP, Lambda, Modal, or Slurm.
- Apache-2.0 licensed and fully open source, with no vendor lock-in.
- Oumi is explicitly in beta, and the README warns that some advanced features may change as the platform evolves.
- Full capability needs GPU hardware and a heavy Python/PyTorch environment, so the barrier to entry is higher than with lightweight tools.
- Cloud job submission and dedicated endpoint deployment depend on external accounts and quotas (AWS, Azure, GCP, Lambda, Modal, Slurm, Fireworks, Parasail).
- Many supported model weights ship under custom licenses (Llama, Gemma, Qwen and others), so commercial use requires checking each license.
- The README's news list contains dates that run into 2026, so release history should be verified against the actual release pages.
How does this agent compare with similar options?
Key facts side by side with the most closely related agents.
| Agent | Source review | Form / cost | Stars | Updated | Language | Full support on |
|---|---|---|---|---|---|---|
| Oumi This agent | 58 · Major gaps | CLIFree + model costs | ★ 9.4k | today | Python | Claude Code · OpenAI API |
| AI Agents Projects & Tutorials | 9 · Major gaps | Library / SDKFree + model costs | ★ 2.9k | 2d ago | Jupyter Notebook | OpenAI API · Claude API |
| MARTI: Multi-Agent Reinforced Training & Inference for LLMs | 36 · Major gaps | CLIFree | ★ 560 | 1mo ago | Python | — |
| AgentsMeetRL — Awesome List of Agentic Reinforcement Learning | 29 · Major gaps | Agent plugin / skillFree | ★ 1.9k | 16d ago | HTML | — |
How does FollowAgents rate this agent?
Why each dimension lost points
Evidence is limited to README, LICENSE, SECURITY.md, pyproject.toml and a few CI/test files. least_privilege: README shows `oumi launch up` creating cloud resources and `oumi deploy` deploying endpoints, but never states required minimal permissions or IAM scope, so 1. user_confirmation: no confirmation or dry-run described for training, deployment or cloud launch, so 1. data_flow_transparency: README lists a posthog telemetry dependency and tests set DO_NOT_TRACK, but README does not explain data flows, so 1. sensitive_data_handling: no guidance on secrets, credentials or PII, so 1. dependency_security: pyproject bounds most dependencies with explanatory comments and a CodeQL workflow exists, which is moderate, so 2. external_effects: side effects of training, deployment and cloud provisioning are not described with scope or recovery, so 1. rollback: no rollback or recovery procedure documented, so 1. source_attribution: full Apache-2.0 text with copyright holder and a SECURITY.md disclosure channel, so 2.
self_consistency: README install, CLI and recipe paths align with pyproject entry points (oumi, oumi-mcp), but the README is truncated so all links cannot be checked, so 2. dependency_availability: dependencies come from PyPI with version ranges and comments explaining the omegaconf dev pin and lm_eval cap, so 2. failure_messages: README mentions partial-failure support but gives no error text or diagnostics examples, so 1.
audience_and_scenarios: README explicitly targets researchers, enterprises and laptop-to-cluster-to-cloud scenarios, so 3. capability_boundaries: supported models and training methods are listed, but unsupported capabilities and boundaries are not systematically stated, so 2. trigger_precision: as a framework/CLI there is no trigger condition or invocation-timing guidance, so 1. environment_fit: covers pip, Docker, install script, multiple Python versions and many GPU/cloud platforms, so 3.
information_architecture: README is well structured (News/About/Getting Started/Usage/Examples) but the file is truncated, so 2. install_notes: pip, uv, Docker, experimental install script and GPU variants are well documented, so 3. naming_stability: CLI commands and config paths are consistent, but no stability commitment is made, so 2. examples_and_faq: many Colab notebooks and recipe tables, so 3. known_limitations: only a one-line beta note, no systematic limitations, so 1. license: full Apache-2.0 text, so 3. versioning_changelog: the News section lists multiple releases with links, so 3. maintenance_responsibility: SECURITY.md describes team response process, but publisher identity is unverified, so 2.
output_usability: CLI and recipes produce training/eval/inference outputs directly, but output formats are not described, so 2. marginal_value: integrates training, evaluation, deployment and data synthesis, adding value over scattered tooling, so 2. cost_benefit: very heavy dependency set (torch, vllm, verl, deepspeed, etc.) implies high install and resource cost, and README does not discuss trade-offs, so 1.
claim_traceability: README claims SOTA, enterprise-grade and production-grade reliability without benchmarks or evidence links, so 1. cross_source_corroboration: README and pyproject corroborate dependencies and entry points, but safety and reliability claims have no second source, so 1. fact_inference_separation: README mixes marketing language with facts and does not separate claims from evidence, so 1.
- README claims 'production-grade reliability' and 'enterprise-grade' without benchmarks, test reports or verifiable evidence; treat as marketing language.
- Cloud provisioning (oumi launch up) and endpoint deployment (oumi deploy) lack least-privilege guidance, confirmation steps and rollback instructions; assess permissions and cost before use.
- The dependency set is very heavy (torch, vllm, verl, deepspeed, etc.), raising install and runtime cost, and some dependencies use dev versions or tight caps that may cause conflicts.
- Publisher identity is unverified; maintenance responsibility and update path can only be inferred from SECURITY.md and release notes inside the repository.
- The README is truncated in the evidence, so some links and claims cannot be fully checked.