July 2026 Summaries
4 posts from Blacksmith
Filter
Month:
Year:
Post Summaries
Back to Blog
Docker build performance depends on three distinct caches—layer caches, persistent mount caches for package managers and compilers, and the Docker image store—while changes trigger an invalidation wave that reruns affected layers and those after them. The discussion argues that teams should optimize for common source and dependency changes rather than only cold or unchanged rebuilds, particularly by ordering Dockerfiles so stable, expensive dependency-install steps precede frequently changing source files. Benchmark results across several languages indicate that mount caches can be crucial for incremental compiled builds but are generally not preserved by portable layer-cache exporters such as GitHub Actions or registry caches, whose import and export overhead can sometimes make changed builds slower than no caching. Persistent builder state can retain both layers and mounts between jobs, though it must be scoped by shared image lineage to avoid eviction among unrelated images or duplicated shared bases in monorepos. Cache size is driven chiefly by distinct dependency versions, with cleanup behavior varying by toolchain, while separate organization-wide image-store caching can reduce repeated pulls of common large images and improve fan-out workloads after an image has entered the shared store. The central recommendations are to match caching mechanisms to where reuse occurs, weigh transfer costs against cache hits, share state only among builds with common layers, and avoid unnecessary image serialization during distribution.
Jul 31, 2026
2,838 words in the original blog post.
Blacksmith’s CI platform treats storage as a central systems challenge because it creates and destroys roughly 2.5 million ephemeral per-job volumes daily, with peaks near 1,000 new microVMs per minute, generating about two petabytes of writes per day that are mostly discarded. Each short-lived job must quickly boot from shared images, fetch repositories and dependencies, rebuild environments, run I/O-intensive tests, preserve useful caches across VM lifetimes, and durably retain artifacts such as logs and build outputs. The workload combines extreme churn, high concurrent access to shared data, competing cache updates, strict in-run reliability, and differing durability requirements, since most data can disappear after a run while artifacts must remain available long term. Conventional block devices, local tool caches, and object stores address pieces of the problem but do not inherently provide CI-specific capabilities such as lazy data access, incremental updates, cache consistency, policy-aware eviction, locality, and protection against corrupted shared state. The author argues that pooling many customers’ bursty workloads can improve infrastructure utilization and identifies open systems problems around fast data delivery, deduplicated warm environments, online classification of write durability, artifact persistence, and safe sharing, describing these challenges as the reason for joining Blacksmith’s storage and systems engineering effort.
Jul 24, 2026
3,676 words in the original blog post.
On July 21, 2026, Blacksmith’s control plane suffered major degradation from about 10:20 AM to 3:55 PM ET after a burst of erroneous, largely duplicate GitHub webhook events overwhelmed an already heavily loaded Redis instance used to process and queue customer jobs. Redis CPU saturation slowed requests, exhausted HTTP workers, caused cache, monitor, and sticky-disk operations to fail or time out, slowed job execution, and created a substantial backlog of unstarted jobs; most regions returned to normal queueing by 5:25 PM ET, eu-central recovered around 7:00 PM ET, and failed dispatches were reissued by 7:19 PM ET. Recovery was prolonged when an attempted HTTP worker scale-up exhausted database connections and shared-pool health checks caused the load balancer to terminate healthy instances, obscuring the Redis root cause and reducing available capacity. Blacksmith resolved the incident by distributing Redis traffic across multiple instances and moving load-balancer checks to an unaffected status endpoint, and it plans to strengthen independent health checks, monitoring, alerting, capacity targets, and broader control-plane architecture. The company also acknowledged inadequate customer communication and committed to a formal incident process with a communications owner and status updates at least every 30 minutes.
Jul 22, 2026
1,038 words in the original blog post.
Blacksmith describes a host-based network observability system for ephemeral, untrusted CI virtual machines that records outbound traffic by job and, where possible, by domain name to diagnose slow builds and support future egress controls. Because guest workloads may be hostile and must remain unchanged, the design operates entirely outside each VM at its host-side virtual network interface, using iptables only to redirect DNS traffic to a host proxy and eBPF programs attached through Linux TCX to count traffic in both directions with minimal overhead. The DNS proxy associates observed DNS answers with destination IPs, while isolated per-VM eBPF maps maintain cumulative counters for each IP, port, and protocol combination. Rather than generating per-packet events or performing many userspace map lookups, BPF iterator programs export compact map snapshots once per second, which are stored in ClickHouse with job, VM, destination, timing, and traffic data that can be correlated with CI steps. The system is designed with bounded resources and visible overflow reporting, but domain attribution remains best-effort because shared IP addresses, hard-coded addresses, cached DNS, and encrypted DNS can prevent definitive name mapping; it currently supports IPv4 only. A planned second phase will use the same architecture to enforce domain-based default-deny egress policies in the host kernel.
Jul 20, 2026
2,630 words in the original blog post.