OMA Skills Hub by OMA · Open Modus Accelerator
🎯 Use Case Hermes Skill

slime-rl-training

RL post-training for LLMs with Megatron and SGLang.

仓库: NousResearch/hermes-agent · v1.0.0 · by Orchestra Research · MIT

Reinforcement LearningMegatron-LMSGLangGRPOPost-TrainingGLM
文件路径
optional-skills/mlops/slime/SKILL.md

查看源文件 (raw) ↗