Agent post-training

Agents that get better from their own failures.

We read your agent's trajectories, find where it fails, build the tasks that target those failures, and train against them. Held-out evals show whether the score really went up.

Approach

Fix the cases that actually fail.

Average traffic is already handled. Gains come from the rare, hard cases, so we start from your real failures and make more of them.

01

Failure modes, not vibes

Trajectories are clustered into named failure modes you can read, dispute and prioritise.

02

Tasks with verifiable rubrics

Each mode becomes new tasks with graded rubrics, reviewed by hand before anything trains on them.

03

Honest evaluation

Tasks are split into held-out evals and training tasks up front, so the reported gain is on cases the model never trained on.

04

Harness-side skills

Beyond weights, we add auditing skills to the agent harness, which raised performance further.

How it works

Five steps, run by us.

  1. Collect trajectories

    You share agent runs, with outcomes if you have them.

  2. Diagnose failure modes

    We identify and rank how the agent fails.

  3. Generate eval and train tasks

    New tasks per failure mode, split into eval and train.

  4. Train with RL

    Our RL pipeline runs on our infrastructure. You do not manage GPUs.

  5. Verify on held-out

    You receive the eval report, the tasks and the improved agent.

Case study

A production spreadsheet agent.

We ran the full loop on a deployed agent that edits and audits financial workbooks: failure modes from real trajectories, new tasks, an eval/train split, RL, and a harness-side auditing skill.

45%held-out score, before
55%after RL on generated tasks
72.4%plus auditing skill in harness
1,000+tasks generated

Scores on our in-house auditing benchmark. The base model with the auditing skill alone scores 50.6%.

Fully managed

You send trajectories. You get a better agent.

Data pipeline, task generation, evals, RL training and versioning are ours.

curl https://api.upshiftlabs.llc/v1/chat/completions \
  -H "Authorization: Bearer $UPSHIFT_KEY" \
  -d '{"model":"your-agent-v2","messages":[...]}'

Contact

Tell us what your agent gets wrong.

Or write to hello@upshiftlabs.llc