Home / Companies / Qdrant / Blog / Post Details
Content Deep Dive

Qdrant and Minima Deliver 2.92x More Agentic RAG Tasks per GPU-Hour

Blog post from Qdrant

Post Details
Company
Date Published
Author
Qdrant and Minima Engineering
Word Count
1,609
Company Posts That Month
14
Language
English
Hacker News Points
-
Post removed?
No
Summary

Qdrant and Minima report that combining hybrid, reranked retrieval with compressed LLM inference increased agentic RAG performance to 3,158 successful tasks per GPU-hour, or 2.92 times the baseline dense-retrieval BF16 setup, while reducing median task latency from 21.3 to 7.7 seconds and maintaining similar grounded-answer and citation quality. Across 1,800 benchmark tasks from SciFact, FiQA, HotpotQA, and a payload-filtering dataset, Qdrant’s dense-plus-BM25 retrieval, reciprocal-rank fusion, late-interaction reranking, and payload filters made first-pass evidence sufficient in 87% of tasks, reduced retrieved context by 56%, and recorded no tenant-policy violations in 50,000 adversarial filtering queries. Minima ran Qwen3.6-27B on one NVIDIA RTX PRO 6000 Blackwell GPU using NVFP4 weights, reduced-precision KV-cache tiers, and native kernels, shrinking model-weight memory from 54.0 GB to 16.9 GB and nearly doubling standalone inference throughput. The reported GPU-only cost per 1,000 successful tasks declined from $1.39 to $0.48, although the comparison excluded Qdrant hosting and embedding-service costs.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
RAG 8 1,152 209 75 -6%
Vector Search 5 2,358 371 127 +5%
LLM 2 5,068 1,020 229 -34%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.