Serverless Multi-turn RL Fine-tuning for LLM Agents Available on Amazon SageMaker AI
Amazon SageMaker AI introduces a serverless service to fine-tune large language model (LLM) agents using multi-turn reinforcement learning (MTRL). Developers can now train agents for complex search and tool-using tasks with production-scale RL at per-token pricing, no GPU management required.
What changed?
Amazon SageMaker AI now supports multi-turn reinforcement learning (MTRL) as a serverless service for fine-tuning LLMs, specifically for agentic search and tool-using behaviors. MTRL allows you to optimize agents over full multi-step interactions using custom reward functions and modular agent-environment setups. The service is serverless, so there’s no need to provision or manage GPU clusters, and pricing is per token. Training jobs are resumable and metrics are observable via MLflow. PPO and other major RL algorithms are offered out-of-the-box.

Why does it matter to an everyday developer?
This release eliminates much of the operational and algorithmic complexity in productionizing RL-based agents. For use cases like search agents, RAG with tool use, or task automation, you no longer need deep RL expertise or GPU fleet management. Fine-tuned smaller models can achieve reliability close to much larger models, but with lower latency and cost. Developers can define custom tool interactions and conversation shapes directly, using their own reward functions. Results show significant improvement in both retrieval quality and drastic reduction in agent failure rates (e.g., from 22.89% to 0.68% on a benchmark).
What can the developer do now?
- Use Amazon SageMaker AI MTRL to fine-tune custom LLM-powered agents for multi-turn search or tool-using tasks.
- Start with a supported model (such as Qwen3.6-27B), your environment, and tool endpoints (BM25, vector search, etc.).
- Prepare training data (multi-turn interactions) in the required format as outlined in the documentation.
