Home / Companies / Braintrust / Blog / Post Details
Content Deep Dive

Testing whether language model harnesses transfer the wrong strategy

Blog post from Braintrust

Post Details
Company
Date Published
Author
Braintrust Team
Word Count
5,964
Company Posts That Month
7
Language
English
Hacker News Points
-
Post removed?
No
Summary

An experiment tested whether a Recursive Language Model harness trained to decompose tasks into chunks might transfer an incorrect “sum the chunk results” strategy to a superficially similar task where global deduplication is required. Using Qwen3-30B with and without an RLM LoRA adapter, the study compared COUNT, where chunk totals can correctly be added, with SENDERS, where summing distinct senders per chunk creates double-counting. Although the task design controlled for reading difficulty, pseudonymized names, and alternative overcounting strategies, repeated evaluations encountered floor and ceiling effects: an initial version was too difficult, while a redesigned version made SENDERS too easy. A pooled analysis of 13 pairs produced an apparent negative transfer result under one random seed but not another, with the estimated accuracy effect shifting by roughly 0.4, while both models double-counted at identical rates in every run. The author concludes that the evidence neither demonstrates nor rules out harmful strategy transfer, partly because the experiment did not reliably measure whether the harness treated the two tasks as equivalent, and recommends future tests using identical sender records arranged in layouts where chunk-level addition is either valid or invalid.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 5 1,189 251 109 -83%
AI Model Fine-tuning 2 103 37 26 -89%
Observability 2 625 152 84 -84%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.