October 2026 Summaries
6 posts from LaunchDarkly
Filter
Month:
Year:
Post Summaries
Back to Blog
AI agents often remain constrained by human review and slow approval processes because their behavior can change unpredictably with updates to models, prompts, tools, or data sources, making traditional deterministic software practices insufficient. Conventional operational metrics such as cost, latency, token use, and tool failures do not measure whether an agent’s output is substantively correct, so teams are encouraged to define quality criteria, use model-based evaluation on production traffic, and focus human review on flagged cases. When quality declines, manual investigation and deployment cycles can leave users exposed to failures, whereas predefined automated responses such as reverting to a known-good configuration, using a simpler fallback, or escalating to a human can limit risk quickly. Continuous improvement is also hindered when every modification requires lengthy testing and approval, so teams can instead evaluate changes against real production examples, gradually test them alongside existing versions, and automatically roll them back if quality falls. The broader argument is that organizations need processes that can measure output quality, respond rapidly to degradation, maintain records of changes, and make small reversible improvements in order to deploy more capable and trustworthy agents.
Oct 03, 2026
1,272 words in the original blog post.
Feature engineering is presented as a central determinant of production machine-learning performance, transforming raw data into meaningful, reliable inputs through derived behavioral variables, categorical encodings, numerical scaling, temporal aggregates, and suitable representations for deep learning and text. Effective practice begins with understanding data provenance, semantics, availability, and the data-generating process to prevent leakage, bias, missing-data errors, and unstable correlations. The material emphasizes implementing preprocessing as versioned, reproducible code; fitting parameters only on training data; maintaining point-in-time correctness for temporal features; and using identical transformations in training and serving to avoid skew. It also discusses selecting features according to model and operational constraints, monitoring feature quality and drift, and using feature stores to centralize definitions, lineage, governance, and consistent offline and online access. Finally, it describes feature flags and runtime configuration tools such as LaunchDarkly as mechanisms for gradually testing inference-time changes, including feature sets, model variants, prompts, and LLM-based extraction schemas, while logging configurations to preserve reproducibility and enable safe rollback.
Oct 03, 2026
3,984 words in the original blog post.
AI governance is presented as an ongoing operational discipline for keeping production AI systems safe, reliable, and accountable as user behavior, data sources, models, and dependencies change over time. Its seven connected pillars are fairness and bias mitigation, implementable policies and procedures, data governance, privacy and data protection, risk management, monitoring and evaluation, and runtime configuration management. The approach emphasizes defining measurable fairness outcomes, testing and monitoring performance across user cohorts, restricting and tracing approved data sources, minimizing sensitive-data exposure, and enforcing authorization through application controls rather than relying on the model. It also recommends identifying failure modes, applying layered safeguards, preparing restricted fallback modes, and using offline tests alongside live quality, safety, cost, and reliability signals to guide staged rollouts. Versioned prompts, model settings, tool permissions, and routing rules—with audit trails, controlled experiments, and rapid rollback capabilities—are described as central mechanisms for connecting governance requirements to day-to-day production operations and supporting compliance with external standards and regulations.
Oct 03, 2026
4,126 words in the original blog post.
Managing AI features that rely on hosted models requires a lifecycle distinct from conventional code deployment because teams must frequently adjust model selection, prompts, parameters, targeting, and safeguards without redeploying applications. The described approach centers on externalizing these decisions into versioned runtime configurations, evaluating proposed variations against representative and adversarial datasets with calibrated LLM-based judges, and releasing changes gradually to targeted audiences while monitoring quality, cost, latency, token usage, and errors. Guarded rollouts and automatic rollback can limit the effects of regressions, while production evaluation samples live traffic to identify issues not captured in pre-production testing. Experimentation complements quality evaluation by measuring whether model changes improve business outcomes such as conversion or task completion, with consistent variation assignment needed for reliable attribution. The article presents LaunchDarkly AgentControl as an integrated implementation of these capabilities, combining managed model and prompt configurations, playground-based evaluation, audience targeting, rollout controls, monitoring, adaptive fallbacks, and experimentation, while emphasizing that production failures should become future test cases in a continuous feedback loop.
Oct 03, 2026
3,192 words in the original blog post.
A tutorial describes how to build a small AI-assisted software factory in GitHub using seven stations that move work from a structured issue to verified production: intake, context, planning, execution, review and policy, delivery, and observability. The approach uses issue templates to define goals, constraints, and completion criteria; a maintained CLAUDE.md file to provide repository context; read-only agents to propose plans for human approval; and coding agents with limited permissions to implement approved work and open pull requests. Independent safeguards include rerunning tests, linting, and type checks outside the agent, enforcing branch protections and policies that prevent agents from changing factory configuration, and using a separate reviewer agent to inspect code and test coverage. After human-approved merges, smoke tests validate the deployed application and can automatically create a revert pull request if failures occur, while prompts, plans, and transcripts are retained as artifacts for traceability. Using a URL-shortening application as an example, the tutorial reports that a change adding 30-day link expiration moved from planning through verified deployment in minutes, while emphasizing that the most important part of an AI software factory is its ability to reject unsafe or incorrect changes at every stage.
Oct 03, 2026
2,457 words in the original blog post.
A robust production machine learning architecture separates a data layer, which creates and versions validated inputs, features, and AI-generated outputs, from a runtime layer that controls deployment of trained model artifacts through feature flags. The model artifact connects these layers by recording the exact datasets, feature definitions, code, environments, prompts, models, parameters, and tool configurations used during training, ensuring reproducibility and preventing training-serving skew. AI-generated features such as summaries, classifications, and embeddings must be treated like conventional feature transformations: tested offline, version-locked, pinned in training manifests, and never changed for a deployed downstream model without retraining. The pipeline should validate both raw and transformed data, preserve point-in-time correctness, maintain consistent offline and online feature definitions, and use artifact-based handoffs to support auditing, debugging, and reruns. Candidate models should be promoted through staged, guarded rollouts that shift traffic gradually while monitoring model quality, operational reliability, AI-step metrics, and business outcomes, with feature flags serving as the single runtime release control and enabling rapid rollback. Continuous monitoring should capture model and AI-feature versions, predictions, relevant inputs, system health, drift, and eventual outcomes so teams can distinguish regressions caused by data, features, AI configurations, or model versions.
Oct 03, 2026
4,288 words in the original blog post.