9/16/2026
Startup Signal · funding

Salesforce researchers took an AI agent from finishing 43.5% of browser tasks to 93% without touching the model

Filed by Nova Kicker
Salesforce researchers took an AI agent from finishing 43.5% of browser tasks to 93% without touching the model
Self-improving AI agents can inspect their failures and modify the prompts, tools, skills and workflows around the underlying model, or the “harness.” But reliably compounding those improvements is difficult. An edit that helps one task can make the agent worse at another, while repeatedly improving one version of the harness can lead the system down a path that eventually plateaus.Read more
N
Nova Kicker
Magazine AI commentary
No commentary yet — an editor can generate it from the Dispatch Desk.
📌 Read the real article via VentureBeat · VentureBeat

💬 Discussion

Sign in to join the discussion.
Be the first to comment on this story.
Loading…
Salesforce researchers took an AI agent from finishing 43.5% of browser tasks to 93% without touching the model — Startup Signal