Home / Companies / Hugging Face / Hacker News

Hugging Face on HN

135 posts with 25+ points since 2022

Filters
Since:
Posts by Month (135 total)
Hacker News Posts
Title Points Comments Date
Qwen 3.8 27B 1,437 792 2026-08-14
Kimi-K3 on HuggingFace 1,382 544 2026-07-27
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf] 982 465 2025-12-01
Qwen3.8-2.4T 712 171 2026-08-12
Uncensor any LLM with abliteration 586 287 2024-06-13
Show HN: Sweep, Open-weights 1.5B model for next-edit autocomplete 534 153 2026-01-21
Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July … 469 259 2026-07-28
Kimi K2.7-Code: open-source coding model with better token efficiency 463 240 2026-06-12
Deepseek R1-0528 451 250 2025-05-28
LLM Embeddings Explained: A Visual and Intuitive Guide 451 91 2025-07-28
Llama-3.3-70B-Instruct 425 219 2024-12-06
Try Stable Diffusion's Img2Img Mode 415 156 2022-08-29
Show HN: Hacker News archive (47M+ items, 11.6GB) as Parquet, updated every … 408 167 2026-03-14
Open-R1: an open reproduction of DeepSeek-R1 394 234 2025-01-28
Smollm3: Smol, multilingual, long-context reasoner LLM 388 79 2025-07-08
GLM-4.7-Flash 378 135 2026-01-19
Nanonets-OCR-s – OCR model that transforms documents into structured markdown 361 78 2025-06-16
A Replacement for BERT 348 75 2024-12-19
MonadGPT – What would have happened if ChatGPT was invented in the … 323 116 2023-11-24
Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS 323 61 2025-09-02
The Smol Training Playbook: The Secrets to Building World-Class LLMs 265 19 2025-10-30
DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning 264 88 2025-12-01
LLM in a Flash: Efficient LLM Inference with Limited Memory 252 53 2023-12-20
Microsoft Phi-2 model changes licence to MIT 240 90 2024-01-06
Falcon 180B 238 208 2023-09-06
OpenLLaMA 13B Released 229 107 2023-06-18
Kokoro WebGPU: Real-time text-to-speech 100% locally in the browser 227 53 2025-02-07
Hugging Face Releases Agents 214 125 2023-05-10
Inflect-Micro-v2: complete voice in 9.36M parameters 214 29 2026-07-26
Qwen3-4B-Thinking-2507 198 61 2025-08-06
Space secrets leak disclosure 197 83 2024-06-01
BigCode Project Releases StarCoder: A 15B Code LLM 185 37 2023-05-04
Best 7B LLM on leaderboards made by an amateur following a medium … 181 49 2024-01-05
Stability.ai sent a take down request to Runway ML's SD v1.5 citing … 179 72 2022-10-20
We raised $100M for open and collaborative machine learning 175 47 2022-05-09
LFM2.5 2.6B model competitive with 4x larger models 169 40 2026-08-04
Llama 3 8B is almost as good as Wizard 2 8x22B 168 111 2024-04-19
SantaCoder: A new 1.1B code model for generation and infilling 168 74 2022-12-22
Nvidia releases NVLM 1.0 72B open weight model 167 55 2024-10-02
StackLlama: A hands-on guide to train LlaMa with RLHF 165 38 2023-04-06
Explaining the SDXL Latent Space 163 33 2024-02-05
BLOOM: The largest open multilingual language model 160 46 2022-07-12
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence 159 19 2026-04-24
Show HN: Text-to-video model from scratch (2 brothers, 2 years, 2B params) 158 24 2026-01-22
Qwen3-235B-A22B-Thinking-2507 155 64 2025-07-25
Hugging Face and Google partner for AI collaboration 152 57 2024-01-25
Show HN: Penny-1.7B Irish Penny Journal style transfer 149 None 2025-06-02
Wordalle – Guess the prompt used to generate a set of images … 137 None 2022-07-01
Mistral-8x7B-Chat 131 None 2023-12-10
A CC-By Open-Source TTS Model with Voice Cloning 131 None 2024-11-04
Qwen-Image-Layered: transparency and layer aware open diffusion model 130 None 2025-12-19
FineWeb: Decanting the web for the finest text data at scale 127 None 2024-06-02
Yi-34B-Chat 115 None 2023-11-24
GPT-3.5 and Wolfram Alpha via LangChain 107 None 2023-01-18
The Falcon has landed in the Hugging Face ecosystem 105 None 2023-06-05
HuggingChat: Chat with Open Source Models 103 None 2024-02-21
Hugging Face and AWS partner to make AI more accessible 102 None 2023-02-21
HuggingFace Training Cluster as a Service 101 None 2023-09-05
More than 80 AI models from Qualcomm 95 None 2024-02-28
Segmind Stable Diffusion – A smaller version of Stable Diffusion XL 95 None 2023-10-25
LLaMA-Pro-8B 94 None 2024-01-06
HuggingChat 93 None 2023-04-25
Waypoint-1: Real-Time Interactive Video Diffusion from Overworld 92 None 2026-01-23
Yarn-Mistral-7B-128k 88 None 2023-11-11
Qwen3 30B-A3B 87 None 2025-07-30
Apple/OpenELM: Efficient Open-Source Family Language Models 82 None 2024-04-24
Sparse LLM Inference on CPU: 75% fewer parameters 78 None 2023-10-19
Pokemon GAN 77 None 2022-02-14
YouTube-Commons: Audio transcripts of 2,063,066 YouTube videos, CC-By license 75 None 2024-04-18
Switch Transformers C – 2048 experts (1.6T params for 3.1 TB) (2022) 73 None 2023-11-20
Multimodal Neurons in Pretrained Text-Only Transformers 66 None 2023-08-04
Show HN: Simply Reading Analog Gauges – GPT4, CogVLM Can't 66 None 2024-01-22
Voxtral-Mini-3B-2507 – Open source speech understanding model 64 None 2025-07-15
Open-sourcing 5,000hrs of self-driving dataset 63 None 2025-03-11
DeepSeek-v3.2 63 None 2025-12-01
HuggingChat – ChatGPT alternative with open source models 61 None 2023-12-15
MSFT's WizardLM2 models have been taken down 58 None 2024-04-16
OpenLLaMA 7B Training Completed to 1T Tokens 58 None 2023-06-07
Phi-2 57 None 2023-12-13
Dolphin-2_6-Phi-2 56 None 2023-12-24
Alibaba releases 72B LLM with 32k context length 55 None 2023-11-30
Show HN: 30k IKEA items in flat text 55 None 2026-01-07
LiteLlama-460M-1T has 460M parameters trained with 1T tokens 54 None 2024-01-07
Qwen Image 54 None 2025-08-04
Fine-Tuning LLMs to 1.58bit 52 None 2024-09-18
Train faster static embedding models with sentence transformers 52 None 2025-01-15
Show HN: ChatToSTL – AI text-to-CAD for 3D printing 52 None 2025-06-12
LLaMA 3 70B Llamafiles 51 None 2024-04-19
Janus-Pro: Autoregressive framework unifying multimodal understanding&generation 49 None 2025-01-27
DeepSeek v3 beats Claude sonnet 3.5 and way cheaper 48 None 2024-12-26
Improving Parquet Dedupe on Hugging Face Hub 47 None 2024-10-08
Open LLAMA 13B released, trained on 1T tokens 47 None 2023-06-19
DALL·E Mini 46 None 2022-04-11
Open-LLM performances are plateauing 46 None 2024-06-29
The AI Research Residency Program 46 None 2022-03-23
Anyone Can Clone Your Voice Now 45 None 2026-01-26
Drax: Speech Recognition with Discrete Flow Matching 45 None 2025-11-09
Show HN: Chonky – a neural text semantic chunking goes multilingual 43 None 2025-10-25
4-Bit Quantization and QLoRA 41 None 2023-05-25
Qwen/Qwen3.6-27B · Hugging Face 41 None 2026-04-22
BLOOMChat, a 176B parameter, Multi-lingual, fine tuned chat 40 None 2023-05-19
What's Going on with the Open LLM Leaderboard? 40 None 2023-06-23
Kai-Fu Li's Yi-34B uses exactly Llama's architecture except for 2 tensor renamed 39 None 2023-11-14
DeepSeek-R1-Distill-Qwen-1.5B Surpasses GPT-4o in certain benchmarks 39 None 2025-01-20
Continuous batching (2025) 39 None 2026-02-15
Fully autonomous AI agents should not be developed 38 None 2025-02-07
Zephyr 7B – Mistral Finetune that responds like ChatGPT 37 None 2023-10-15
Whisper Jax: Transcribe a 1 hour of audio in under 15 seconds 36 None 2023-04-22
Qwen3-235B-A22B-Instruct-2507 36 None 2025-07-21
MistralLite by Amazon Web Services 34 None 2023-11-01
Mixtral-8x22B on HuggingFace 33 None 2024-04-10
The Ultra-Scale Playbook: Training LLMs on GPU Clusters 33 None 2025-02-19
Qwen3-Coder-30B-A3B-Instruct 32 None 2025-07-31
General OCR Theory: Towards OCR-2.0 via a Unified End-to-End Model 31 None 2024-09-11
Anatomy of BoltzGen 31 None 2026-01-04
Zephyr 141B, a Mixtral 8x22B fine-tune, is now available in Hugging Chat 30 None 2024-04-12
OpenFLUX.1 30 None 2024-10-04
Reachy Mini – The Open-Source Robot for Today's and Tomorrow's AI Builders 30 None 2025-07-09
Mistral 7B v0.2 29 None 2024-03-31
Mixture of Experts Explained 29 None 2023-12-11
TinyLlama at 2T of 3T 29 None 2023-11-19
Video2Game: Real-Time, Interactive, Realistic Environment from a Single Video 28 None 2024-04-16
Real-Time Latent Consistency Model 27 None 2023-10-30
Language Modeling Is Compression 27 None 2023-09-21
grok-2 on Hugging Face 27 None 2025-08-23
Llama-3.2-3B-Instruct-uncensored 26 None 2024-09-27
Pixel Art XL: Stable Diffusion XL for Pixel Art 26 None 2023-08-03
UC Berkeley's open-source Vicuna LLM chatbot released new improved model weights 26 None 2023-04-14
Llama can now see and run on your device – welcome Llama … 26 None 2024-09-25
DeepSeek-v3.1 26 None 2025-08-21
DeepSeek-V4 Technical Report [pdf] 26 None 2026-04-24
Llama 1.3B Trained on 200B Tokens for Commercial Use 25 None 2023-04-28
New Phi-3.5 Models from Microsoft, including new MoE 25 None 2024-08-20
LLM: Transformer Is Linear 25 None 2024-05-24
DeepSeek-v3.1-Base 25 None 2025-08-19