
post-training
community6,242 stars
RLHF and preference alignment including TRL, GRPO, OpenRLHF, SimPO, verl, slime, miles, and torchforge. Use when aligning models with human preferences, training reward models, or large-scale RL training.
Install
/plugin install post-training@orchestra-research-ai-research-skillsFeatured
Make coding agent sessions - Searchable, Shareable, Vendor-neutral & Scored.
Try For Free →Your agent targets a perfect 10 Code Health score. Deterministic. Every commit.
Try For Free →Integrate web data into your AI product. One API to scrape website & brand data.
Get API Key Now →belt cli automatically finds the best tools and skills for your agent. image, video, music, tts...
one prompt install →Plug Mailtrap into your AI workflow and let it handle the email.
Connect Mailtrap MCP →Agent, run crypto. Access onchain data & trade routes via 1inch.
Install now →