Home / Companies / Baseten / Hacker News

Baseten on HN

40 posts with 1+ points since 2022

Filters
Since:
Posts by Month (40 total)
Hacker News Posts
Title Points Comments Date
Show HN: ChatLLaMA – A ChatGPT style chatbot for Facebook's LLaMA 402 215 2023-03-22
Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs 247 175 2025-08-07
The efficient frontier of LLM inference 155 46 2026-09-01
Show HN: Baseten – Build ML-powered applications 112 11 2022-04-26
DALL-E Mini – Generate images from a text prompt 52 22 2022-06-10
How we got Stable Diffusion XL inference to under 2 seconds 51 5 2023-08-31
Show HN: Free Stable Diffusion 2.0 hosted interface 25 2 2022-11-24
Show HN: Fine-tune generative models in 1 line of code 16 0 2023-03-01
Show HN: Baseten Chains – Framework and SDK for Multi-Model AI Products 9 5 2024-06-27
The Math Behind TurboQuant 8 3 2026-03-27
Hosted Stable Diffusion Demo 7 0 2022-08-24
Try it yourself: Speech to text with Whisper 5 0 2022-10-01
How BaseTen is using “docs as code” 5 0 2022-03-09
Baseten raised a $1.5B Series F and achieved a $13B valuation 5 0 2026-06-22
Serving four million Riffusion requests in two days 5 0 2022-12-21
SDXL inference in under 2 seconds 3 1 2023-08-31
How to get GLM 5.2 to 280 tokens per second 3 1 2026-06-23
How We Built the Fastest Kimi K2.5 on Artificial Analysis 3 0 2026-02-11
Deploying Stable Diffusion in Production Using Truss 3 0 2022-09-01
Faster Mixtral inference with TensorRT-LLM and quantization 2 1 2023-12-27
Show HN: Inference Engineering 2 0 2026-02-23
Open Source Inference Engine Baseten Raises $40M from IVP, Spark and Greylock 2 1 2024-03-14
We built a day-0 API for Kimi K3 2 0 2026-07-27
We built the new fastest API for GLM-5.2 2 0 2026-07-26
Inference Engineering by Philip Kiely – Digital Download 2 0 2026-08-18
Code generation interactive demo (Salesforce Codegen mono 2B) 2 0 2022-07-01
FP8: Efficient model inference with 8-bit floating point numbers 2 0 2024-03-08
Show HN: Automatically Build Nvidia TRT-LLM Engines 2 0 2024-08-01
How to double tokens per second for Llama 3 with Medusa 2 0 2024-08-20
Show HN: 60% higher tokens per second for 70B custom LLMs 1 0 2024-07-31
A guide to LLM inference and performance 1 0 2025-02-16
Introduction to quantizing machine learning models 1 0 2024-02-16
Three techniques to adapt LLMs for any use case 1 0 2023-06-15
Accelerating model deployment: 100X faster dev loops with draft models 1 0 2022-12-09
Demo – Text generation with EleutherAI's GPT-J-6B model 1 0 2022-04-29
Deploying custom ComfyUI workflows as APIs 1 0 2024-11-20
Continuous vs. dynamic batching for AI inference 1 0 2025-08-06
Continual learning and the post monolith AI era 1 0 2026-02-06
Inferless Joins Baseten 1 0 2026-02-16
How to build function calling and JSON mode for open-source and fine-tuned … 1 0 2024-09-12