LLM Alignment and Post-Training
Parent: LLM Models and APIs · researched 2026-05-31T20:37:33.874Z· 18 sources · 11 concepts · skill ai-agent-engineering
AI & agent-engineering family ROUTER. Split into: ai-agents-orchestration (agent frameworks, multi-agent, memory, planning, guardrails, coding/GUI agents, autonomous loops, eval); ai-rag-retrieval (RA
ai-agent-engineering
- AI & agent-engineering family ROUTER. Split into: ai-agents-orchestration (agent frameworks, multi-agent, memory, planning, guardrails, coding/GUI agents, autonomous loops, eval); ai-rag-retrieval (RAG, iterative retrieval, vector/graph datastores); ai-llm-model-layer (training, fine-tuning, alignment/RLHF, compression, inference serving, transformer/multimodal architecture, model selection, observability); ai-mcp-sdk-prompting (MCP servers/builder, Anthropic SDK, prompt engineering, context engineering, LLM frameworks, tool-search, prompt lookup). Route to the matching sub-hub. [source]
- This hub routes to on-demand reference files under references/. See each spoke for depth. [source]
Children
- Supervised Fine-Tuning / Instruction Tuning (frontier)
- Reward Modeling (Bradley-Terry, RewardBench) (frontier)
- RLHF with PPO (KL penalty, value model) (frontier)
- RLAIF (AI feedback) (frontier)
- Constitutional AI / RLCAI (frontier)
- DPO (Direct Preference Optimization) (frontier)
- DPO-Variant Family (IPO/KTO/ORPO/SimPO/CPO) (frontier)
- Preference-Data Pipelines (pairwise/ratings, on/off-policy, iterative/self-rewarding) (frontier)
- Reward Hacking / Length Bias / Over-Optimization (frontier)
- Alignment Evaluation (win-rate, LC-AlpacaEval, Arena-Hard, RewardBench, safety) (frontier)
- Alignment Tooling (TRL, alignment-handbook, Axolotl, OpenRLHF) (frontier)
Frontier under this node: Alignment Evaluation (win-rate, LC-AlpacaEval, Arena-Hard, RewardBench, safety), Alignment Tooling (TRL, alignment-handbook, Axolotl, OpenRLHF), Constitutional AI / RLCAI, DPO (Direct Preference Optimization), DPO-Variant Family (IPO/KTO/ORPO/SimPO/CPO), Preference-Data Pipelines (pairwise/ratings, on/off-policy, iterative/self-rewarding), RLAIF (AI feedback), RLHF with PPO (KL penalty, value model), Reward Hacking / Length Bias / Over-Optimization, Reward Modeling (Bradley-Terry, RewardBench), Supervised Fine-Tuning / Instruction Tuning