Skip to main content
← Back to News
AI Tools

Fine‑tune Search Agents with Multi‑turn RL on SageMaker

AWS details how to fine‑tune a Qwen3.6‑27B search agent using Amazon SageMaker AI multi‑turn reinforcement learning, achieving higher retrieval quality and lower failure rates.

Published: October 2, 2026By GetAISet Editorial

Original source published: October 2, 2026

AWS explains that search agents powered by large language models can autonomously decide what to query, which retrieval strategy to use, and when to stop, but achieving reliable multi‑turn behavior is difficult. The blog introduces Amazon SageMaker AI Multi‑turn Reinforcement Learning (MTRL), which frames an agentic task as a sequence of decisions and optimizes it with policy‑gradient algorithms while keeping integration low‑code and execution serverless.

The authors fine‑tuned a Qwen3.6‑27B model for an enterprise search scenario that accesses BM25 lexical and vector‑search tools. Training used several public and synthetic datasets, a reward based on nDCG@10, and a minimal hyper‑parameter change (max_epochs = 1, batch = 128, rollout concurrency = 32). MTRL’s default settings handle algorithm selection, advantage estimation, and off‑policy bounds.

Evaluation on four held‑out benchmarks showed improvements on three datasets, with the largest gains on BrowseComp‑Plus (+23.7 % nDCG@10) and WixQA (+18.4 %). Failure rates also dropped dramatically, exemplified by BrowseComp‑Plus falling from 22.89 % to 0.68 %, indicating more reliable task completion.