Forge scores every completed run with an AI evaluator. Over time those scores reveal which agent on a team is holding quality back — and Forge can rewrite that agent's prompt for you. This guide walks through the Self-improvement panel end to end.
Step 1: Open the team and find Self-improvement
Go to Teams and open a team you own. The Self-improvement panel shows the team's rolling quality score, a trend arrow (improving or declining), and a per-criterion breakdown — completeness, accuracy, actionability, and clarity. The weakest criterion is highlighted.
3 evaluated runs to learn from. If you just created the team, run it a few times first — the panel tells you how many evaluated runs you have.Step 2: Generate a suggestion
Click Suggest an improvement. Forge reads the rationales from your lowest-scoring runs — and any low-star marketplace reviews — and proposes a rewritten system prompt for the single most impactful agent, with a one-line summary of what changed and why.
Step 3: A/B test before you commit
Click A/B test to run the current prompt against the proposed one on a representative brief. An impartial judge picks a winner and shows its reasoning inline — so you apply changes on evidence, not a hunch.
Step 4: Apply — and revert if needed
Click Apply to swap the agent's prompt. The previous prompt is kept, so the suggestion appears under Applied improvements with a Revert option. Run the team again and watch the quality trend.
Hands-off: turn on Autopilot
Flip the Autopilot toggle at the top of the panel and Forge does all of the above on its own. A daily job suggests an improvement for the weakest agent, A/B tests it, and auto-applies the winner when it clearly beats the current prompt — weaker ideas are left as suggestions for you. Every automatic change notifies you and can be reverted.