Home / Companies / Onehouse / Blog / Post Details
Content Deep Dive

Announcing AI training data pipelines on Onehouse: feeding open data to open models

Blog post from Onehouse

Post Details
Company
Date Published
Author
-
Word Count
4,116
Company Posts That Month
1
Language
English
Hacker News Points
-
Post removed?
No
Summary

Onehouse introduces a lakehouse-based workflow for creating reproducible AI training datasets from operational data, agent traces, feedback, and business outcomes, addressing the limitations of ad hoc exports that lose lineage, version history, and correction paths. Its platform uses OneFlow to ingest data into open Apache Hudi and Iceberg tables, Quanton to curate, validate, version, and export training examples, and optional Apache Airflow orchestration, while Baseten, Fireworks, and Together AI perform the actual model training and Lakegres provides inference-time context. The approach emphasizes that post-training, fine-tuning, distillation, and agent improvement require carefully selected examples tied to explicit evaluations and outcomes rather than unreviewed production logs. It highlights evidence from legal and coding-agent post-training efforts showing that measurable gains depend on well-defined tasks, realistic environments, expert evaluation criteria, and reliable data selection. A contract-analysis example illustrates how labeled table data can be transformed into supervised fine-tuning examples, while noting the need for balanced negative cases, stable source identifiers, validation, and immutable dataset versions. By retaining source snapshots, lineage, training configurations, and provider handoff records, the system aims to let teams inspect past training inputs, identify effects of changed labels or data, rebuild datasets as feedback arrives, and connect the continuous cycle of training, evaluation, inference, and operational outcomes.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Model Fine-tuning 7 139 28 14 -75%
Vector Search 5 265 57 33 -89%
Observability 3 472 102 54 -85%
AI Agents 1 931 231 103 -84%
AI Guardrails 1 35 22 12 -94%
Data Pipeline 1 34 23 18 -90%
Harness engineering 1 33 23 14 -84%
LLM 1 747 162 79 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.