Home / Companies / Neon / Blog / Post Details
Content Deep Dive

Open-weight models are fast on Neon AI Gateway. Here's why

Blog post from Neon

Post Details
Company
Date Published
Author
Carlota Soto
Word Count
1,578
Company Posts That Month
14
Language
English
Hacker News Points
-
Post removed?
No
Summary

Neon AI Gateway serves Databricks-hosted open-weight models through Databricks Foundation Model APIs, using an inference stack designed to improve real-world latency and throughput for workloads such as coding agents and multi-step tool loops. The post attributes performance to continuous batching, paged KV-cache allocation, combined prompt-prefill and token-decoding scheduling, TensorRT-LLM-based runtimes with custom GPU kernels, quantization, and hardware-specific multi-GPU tuning, particularly for Mixture-of-Experts models. It emphasizes prompt caching as especially useful for agents that repeatedly send stable system prompts, tool definitions, and examples, citing a Databricks batch-inference deployment of gpt-oss where a roughly 30% cache-hit rate corresponded with 2.5 times higher per-replica input-token throughput and three times lower median latency. Developers can access the gateway through compatible existing SDKs, pay standard model token rates without a Neon surcharge, and, during the beta period, use free tokens on eligible Neon plans in AWS us-east-2 to evaluate performance on their own workloads.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 10 4,718 960 222 -38%
Loop engineering 2 64 43 35 -56%
Real-time 1 4,120 979 214 -36%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.