Amazon Web Services (AWS) has introduced a new capability within its Amazon SageMaker AI platform designed to significantly improve the performance and reliability of large language model (LLM)-powered search agents. This advancement utilizes multi-turn reinforcement learning (MTRL) to enable search agents to learn complex, multi-step decision-making processes crucial for enterprise information retrieval.
Information was available with The Chenab Times detailing how Amazon SageMaker AI MTRL allows developers to fine-tune LLMs with reinforcement learning in multi-turn interaction scenarios. This approach addresses the inherent difficulties in training agents to effectively navigate tool usage and complex environments across multiple decision points. Traditional methods, such as supervised fine-tuning with expert demonstrations or single-turn reinforcement learning, often fall short in optimizing the sequential nature of agentic tasks. MTRL, conversely, trains the agent to optimize its entire sequence of actions, ensuring that decisions made in earlier turns positively contribute to the final outcome.
Understanding Amazon SageMaker AI MTRL
Amazon SageMaker AI MTRL is engineered to frame agentic tasks as a series of interconnected decisions. It employs multi-turn rollouts to generate training data and subsequently optimizes the model using policy gradient algorithms. The platform offers a modular agent-environment interface, facilitating low-code integration where users can define custom rewards, tool loops, and conversation structures. Its serverless execution model allows for production-scale agentic RL at per-token pricing, eliminating the need for provisioning and managing specialized GPU clusters.
Key features of SageMaker AI MTRL include asynchronous rollout and trajectory collection, which enable parallel generation and gradient updates while maintaining bounded off-policy staleness to ensure training efficiency without policy drift. It also provides a native algorithm library featuring options like Proximal Policy Optimization (PPO), Clipped Importance Sampling Policy Optimization (CISPO), and importance-sampling (IS) losses, paired with group-based advantage estimators. Resumable training capabilities allow for long training runs to be split across multiple jobs, accommodating time limits. Furthermore, SageMaker AI MTRL offers trajectory and reward observability through MLflow managed by Amazon SageMaker AI, allowing users to inspect agent behavior turn by turn. Evaluation jobs are also integrated, enabling the reporting of reward, pass@k, and trajectory metrics before deployment to SageMaker endpoints or Amazon Bedrock.
Application in Enterprise Search Agents
The MTRL framework is particularly well-suited for search agents, where a clear reward signal, such as retrieval quality, can be defined, and a multi-turn interaction loop naturally exists between the agent and its environment. In a practical demonstration, researchers fine-tuned a Qwen3.6-27B model using Amazon SageMaker AI MTRL for an enterprise search agent. This agent was equipped with two primary tools: lexical search (BM25) for exact keyword matching and vector search for semantic and conceptual queries. The objective was to enhance the agent’s ability to autonomously identify, gather, and synthesize information to answer user queries effectively, with a limit imposed on the number of turns to encourage efficiency.
Training Methodology and Datasets
The training process involved several key components, including meticulously prepared datasets, a defined reward function, and the MTRL job configuration. A diverse array of datasets was utilized for training and testing, encompassing benchmarks designed for multi-hop reasoning, retrieval across various domains, enterprise RAG systems, multilingual product search, and complex question-answering scenarios. These included FRAMES, BRIGHT, Enterprise RAG, ESCI, Musique, MLQA for training, and FreshStack, WixQA, BrowseComp-Plus, and Wands for testing. Datasets were preprocessed to meet the format requirements of the MTRL service, with a five percent reservation of training instances for validation.
The primary metric for evaluating the agent’s performance was Normalized Discounted Cumulative Gain at rank 10 (nDCG@10). This metric, a standard in information retrieval, measures the relevance of the top 10 retrieved documents, rewarding systems that place highly relevant items at the top. nDCG@10 was directly used as the trajectory-level reward signal in MTRL, meaning the reward was assigned only after the agent completed its full multi-turn search. To instill desirable behavior, a reward of -1 was assigned to the agent if it reached the maximum number of turns or token budget in a single turn, explicitly training the model to avoid such failure modes.
MTRL Job Configuration and Results
Configuring the MTRL training job proved to be a streamlined process, requiring minimal hyperparameter adjustments. Key parameters such as max_epochs (set to 1), global_batch_size (set to 128), and rollout_max_concurrency (set to 32) were modified, while other advanced RL settings, including the algorithm and advantage estimator, were left at their default values. This simplified configuration process underscores the platform’s accessibility for users without deep expertise in reinforcement learning.
The fine-tuning process demonstrated significant improvements in the search agent’s performance across multiple held-out benchmarks. On the BrowseComp-Plus benchmark, the fine-tuned agent achieved a +23.7 percent gain in nDCG@10, and on WixQA, an +18.4 percent gain was observed. While Wands also showed improvement, FreshStack experienced a slight regression. More critically, the reliability of the agent saw substantial gains, with the failure rate on BrowseComp-Plus plummeting from 22.89 percent to 0.68 percent. This reduction in failures indicates that the agent not only learned to search more effectively but also to complete tasks within its operational constraints, such as turn and token budgets.
Training progress, visualized by the nDCG@10 reward on both training and validation instances, showed a steady rise over training steps, eventually plateauing. This saturation suggests that the model had reached its optimal performance level with the given training duration. Test performance tables further validated these findings, comparing the fine-tuned Qwen3.6-27B model against the original base model across various datasets. The results consistently showed improvements in nDCG@10 scores and a marked reduction in failure rates, reinforcing the efficacy of the MTRL approach for enhancing search agent capabilities.
To conclude, the integration of Amazon SageMaker AI MTRL offers a powerful and accessible method for fine-tuning LLM-powered search agents. By leveraging reinforcement learning across multiple turns, organizations can develop more intelligent, reliable, and cost-effective information retrieval systems, significantly enhancing enterprise productivity and user experience. The platform’s serverless nature, resumable training, and direct optimization against task-specific metrics make it a compelling solution for advanced AI model customization.
The Chenab Times News Desk

