Can open models carry readable silent signals before they speak? Reproducing J-Lens Readouts on Kimi K3 & Qwen3.5-9B
Blog post from Fireworks AI
Researchers applied the Jacobian Lens, a trained probe designed to map transformer hidden states to likely vocabulary, to the open-weight Kimi K3 and Qwen3.5-9B models to examine whether task-relevant concepts can be detected before they appear in generated text. Across paired-copy experiments and ten prompt episodes involving arithmetic, citrus, factual questions, rhyming, sports, code errors, and language tasks, the probe identified relevant token families in internal states even when visible outputs were identical or had not yet named the concepts. A separately fitted lens on Qwen reproduced 19 of 20 vocabulary-family matches originally defined for Kimi, while transcript-region tests found signals during prompt processing, chat-template context, and early response generation. The study notes that results are limited by tokenization, particularly for multi-token numbers such as “21” and “42,” and that fitting the probes required substantial compute. The authors conclude that open models can expose measurable internal representations associated with later outputs, enabling independent investigation of model behavior, though the findings rely on fitted vocabulary probes rather than direct proof of human-like reasoning.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.