Orchestra-Research avatar

post-training

community6,242 stars

RLHF and preference alignment including TRL, GRPO, OpenRLHF, SimPO, verl, slime, miles, and torchforge. Use when aligning models with human preferences, training reward models, or large-scale RL training.

Install

/plugin install post-training@orchestra-research-ai-research-skills
AI writes the code. CodeRabbit catches the slop.
Try For Free
Make coding agent sessions - Searchable, Shareable, Vendor-neutral & Scored.
Try For Free
Your agent targets a perfect 10 Code Health score. Deterministic. Every commit.
Try For Free
Integrate web data into your AI product. One API to scrape website & brand data.
Get API Key Now
belt cli automatically finds the best tools and skills for your agent. image, video, music, tts...
one prompt install
create and run specialised agents in minutes
build now
Plug Mailtrap into your AI workflow and let it handle the email.
Connect Mailtrap MCP
Agent, run crypto. Access onchain data & trade routes via 1inch.
Install now