How we used evals and inference-time compute scaling to generate beautiful QR codes that actually work
Blog post from Modal
Engineers behind qart.codes improved AI-generated artistic QR codes by treating scannability and visual appeal as separate, measurable objectives, prioritizing a 95% scan-rate service-level goal while avoiding aesthetic regressions. Building on ControlNet-guided Stable Diffusion techniques that preserve enough QR-code structure for error correction, they developed automated evaluations using QReader for scan testing and an aesthetic-rating model, then validated those tools against thousands of human judgments. After manually narrowing promising prompts, models, and generation parameters, they ran large-scale offline parameter sweeps and visualized tradeoffs between quality and scannability. When a single generation could not reliably meet the scan-rate target, they applied inference-time compute scaling by generating eight candidates in parallel, evaluating and ranking them by scan success and aesthetics, and displaying the best four. This approach achieved the scan-rate target with under-20-second p95 latency while improving image quality, illustrating how reliable generative-AI applications can be built through iterative eval development, scaled experimentation, and production-time selection rather than relying on compelling but inconsistent demos.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 3 | 4,922 | 763 | 224 | +11% |
| Serverless | 2 | 1,048 | 263 | 99 | +36% |
| RAG | 1 | 1,131 | 232 | 87 | -9% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.