Home / Companies / Paper Compute Company / Blog / October 2026

October 2026 Summaries

2 posts from Paper Compute Company

Filter
Month: Year:
Post Summaries Back to Blog
No summary generated yet.
Oct 07, 2026 1,730 words in the original blog post.
A pipeline built around labeled coding-agent sessions from Paper Compute exports conversations, turns, labels, and training inputs into Databricks Unity Catalog, where they can be inspected and transformed into fine-tuning and evaluation datasets. Labels such as golden, regression, pushback, apology, and no-outcome distinguish reviewed successful sessions, failures, engineer corrections, candidate review points, and unproductive work, enabling traceable selection decisions. Training data consists of eligible or manually approved sessions, while evaluation data includes held-out good sessions and correction-based regression cases, with full sessions kept together to avoid leakage between datasets. Corrections are converted into tests by providing a model only the conversation before the mistaken response and using the later engineer feedback as a separate judging guideline. MLflow tracks datasets, model runs, and comparisons between a base Qwen3-4B coding model and a LoRA-tuned version, using a Databricks-hosted LLM judge to assess whether responses meet recorded expectations. The approach emphasizes reproducible evidence, separation of training from evaluation, inspection of individual outcomes, and careful review and redaction of potentially sensitive exported session data.
Oct 01, 2026 1,825 words in the original blog post.