Home / Companies / Blacksmith / Blog / Post Details
Content Deep Dive

How Blacksmith runs >10M jobs per day

Blog post from Blacksmith

Post Details
Company
Date Published
Author
Andrew Werner
Word Count
2,876
Company Posts That Month
2
Language
English
Hacker News Points
8
Post removed?
No
Summary

Blacksmith redesigned its CI job scheduling system to handle rapid growth in jobs and machines, improve fleet utilization and tenant fairness, and recover more quickly from failures. Its original Redis-based, host-polling model used serialized Lua scripts to prevent duplicate adoption but became a scalability bottleneck, offered limited global policy enforcement, and struggled to place large 32-vCPU workloads without wasting capacity. The new architecture uses a centralized in-memory “assigner” that has a global view of fleet capacity and pushes assignments to hosts, while preserving durability through a queue and agents’ local disks rather than a scheduler database. Multiple assigner instances receive identical agent reports and maintain synchronized fleet views, allowing a leader elected through etcd to make decisions while peers remain hot standbys for rapid failover. The scheduler prioritizes older demand within priority tiers, applies organization-level fairness policies, relies on indexed in-memory structures to meet high placement-rate targets, and is tested using deterministic simulations and production workload traces. To improve placement of wide jobs, it packs smaller jobs tightly at higher utilization to preserve empty hosts and uses targeted drain reservations informed by historical runtime estimates, which reportedly reduced deep-tail latency for 32-vCPU workloads by an order of magnitude.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.