September 2026 Summaries
1 posts from Onehouse
Filter
Month:
Year:
Post Summaries
Back to Blog
Onehouse introduces a lakehouse-based workflow for creating reproducible AI training datasets from operational data, agent traces, feedback, and business outcomes, addressing the limitations of ad hoc exports that lose lineage, version history, and correction paths. Its platform uses OneFlow to ingest data into open Apache Hudi and Iceberg tables, Quanton to curate, validate, version, and export training examples, and optional Apache Airflow orchestration, while Baseten, Fireworks, and Together AI perform the actual model training and Lakegres provides inference-time context. The approach emphasizes that post-training, fine-tuning, distillation, and agent improvement require carefully selected examples tied to explicit evaluations and outcomes rather than unreviewed production logs. It highlights evidence from legal and coding-agent post-training efforts showing that measurable gains depend on well-defined tasks, realistic environments, expert evaluation criteria, and reliable data selection. A contract-analysis example illustrates how labeled table data can be transformed into supervised fine-tuning examples, while noting the need for balanced negative cases, stable source identifiers, validation, and immutable dataset versions. By retaining source snapshots, lineage, training configurations, and provider handoff records, the system aims to let teams inspect past training inputs, identify effects of changed labels or data, rebuild datasets as feedback arrives, and connect the continuous cycle of training, evaluation, inference, and operational outcomes.
Sep 17, 2026
4,116 words in the original blog post.