They Are Still Training on Your Data
Blog post from TrustedRouter
The passage argues that private prompts, code, documents, interactions, and tool traces may be highly valuable sources for AI training even when companies state that they do not train on user data directly. It highlights Google DeepMind’s Generative Data Refinement approach, which rewrites sensitive or toxic real-world examples into “grounded synthetic data” intended to preserve useful structure and diversity while removing risky content, citing tests involving personal information, code repositories, and toxic messages. It contends that data-retention and training commitments can leave room for derivative, anonymized, or de-identified content to be used for model or product improvement, and points to terms from services such as OpenRouter and Vercel as examples of disclosures that may permit such uses under certain conditions. The passage contrasts these concerns with TrustedRouter’s claimed architecture, which uses attested open-source infrastructure and limits router access to prompt and response content, while noting that upstream model providers remain a separate privacy boundary.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.