Home / Companies / Hugging Face / Hacker News

Hugging Face on HN

344 posts with 1+ points in 2025

Filters
Year:
Posts by Month (344 total)
Hacker News Posts
Title Points Comments Date
DeepSeek-v3.2: Pushing the frontier of open large language models [pdf] 982 465 2025-12-01
Deepseek R1-0528 451 250 2025-05-28
LLM Embeddings Explained: A Visual and Intuitive Guide 451 91 2025-07-28
Open-R1: an open reproduction of DeepSeek-R1 394 234 2025-01-28
Smollm3: Smol, multilingual, long-context reasoner LLM 388 79 2025-07-08
Nanonets-OCR-s – OCR model that transforms documents into structured markdown 361 78 2025-06-16
Apertus 70B: Truly Open - Swiss LLM by ETH, EPFL and CSCS 323 61 2025-09-02
The Smol Training Playbook: The Secrets to Building World-Class LLMs 265 19 2025-10-30
DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning 264 88 2025-12-01
Kokoro WebGPU: Real-time text-to-speech 100% locally in the browser 227 53 2025-02-07
Qwen3-4B-Thinking-2507 198 61 2025-08-06
Qwen3-235B-A22B-Thinking-2507 155 64 2025-07-25
Show HN: Penny-1.7B Irish Penny Journal style transfer 149 None 2025-06-02
Qwen-Image-Layered: transparency and layer aware open diffusion model 130 None 2025-12-19
Qwen3 30B-A3B 87 None 2025-07-30
Voxtral-Mini-3B-2507 – Open source speech understanding model 64 None 2025-07-15
Open-sourcing 5,000hrs of self-driving dataset 63 None 2025-03-11
DeepSeek-v3.2 63 None 2025-12-01
Qwen Image 54 None 2025-08-04
Show HN: ChatToSTL – AI text-to-CAD for 3D printing 52 None 2025-06-12
Train faster static embedding models with sentence transformers 52 None 2025-01-15
Janus-Pro: Autoregressive framework unifying multimodal understanding&generation 49 None 2025-01-27
Drax: Speech Recognition with Discrete Flow Matching 45 None 2025-11-09
Show HN: Chonky – a neural text semantic chunking goes multilingual 43 None 2025-10-25
DeepSeek-R1-Distill-Qwen-1.5B Surpasses GPT-4o in certain benchmarks 39 None 2025-01-20
Fully autonomous AI agents should not be developed 38 None 2025-02-07
Qwen3-235B-A22B-Instruct-2507 36 None 2025-07-21
The Ultra-Scale Playbook: Training LLMs on GPU Clusters 33 None 2025-02-19
Qwen3-Coder-30B-A3B-Instruct 32 None 2025-07-31
Reachy Mini – The Open-Source Robot for Today's and Tomorrow's AI Builders 30 None 2025-07-09
grok-2 on Hugging Face 27 None 2025-08-23
DeepSeek-v3.1 26 None 2025-08-21
DeepSeek-v3.1-Base 25 None 2025-08-19
Mistral Small 3.2 (24B-Instruct-2506) 23 None 2025-06-20
DeepSeek-v3.1 23 None 2025-08-19
Open source speech foundation model that runs locally on CPU in real-time 22 None 2025-10-02
Qwen3 235B beats Claude on some code benchmarks 21 None 2025-07-21
Kyutai 1.6B Streaming TTS 21 None 2025-07-03
Selene Mini: Open-sourced SOTA small language-model-as-a-judge 20 None 2025-01-29
DeepSeek-v3.2-Speciale 20 None 2025-12-01
The smallest VLM ever: 250M parameters 19 None 2025-01-23
Show HN: Largest public dataset of electronic circuit files 19 None 2025-12-18
Deepseek V3-0324 18 None 2025-03-24
Supertonic: Ultra-lightweight on-device TTS model open source by Supertone 17 None 2025-11-23
Vector Search with DuckDB 17 None 2025-02-26
DiffuCoder-7B-CpGRPO: A code generation LLM developed by Apple 17 None 2025-07-04
DeepSeek R1 17 None 2025-01-20
Qwen3 0.6B now on HuggingFace (quantized) 16 None 2025-04-28
Sentence Transformers is joining Hugging Face 16 None 2025-10-22
HunyuanOCR by Tencent: A 1B Parameter End to End OCR Expert VLM 15 None 2025-12-03
DeepSeek-R1-0528 performance improvements 14 None 2025-05-29
TeapotLLM- an open-source <1B model for hallucination-resistant Q&A on a CPU 14 None 2025-04-16
DeepSeek-Prover-V2-671B 14 None 2025-04-30
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Compact VLM 13 None 2025-10-16
Nanonets-OCR2-3B – OCR model that transforms documents into structured markdown 13 None 2025-10-14
Co-Doodle with Gemini 13 None 2025-03-19
Open-source DeepResearch – Freeing our search agents 12 None 2025-02-04
FUTO open-sources 1M row keyboard swipe dataset 12 None 2025-04-04
Show HN: The Legal Embedding Benchmark (MLEB) 11 None 2025-10-26
DeepSeek-TNG-R1T2-Chimera 11 None 2025-07-02
smolagents: A simple library to build AI agents 11 None 2025-01-02
Wan2.2-S2V-14B – audio-driven cinematic video generation model 10 None 2025-08-26
Open Source 1.7tb Dataset of What AI Crawlers Are Doing 10 None 2025-07-03
Hugging Face to sell open-source robots thanks to Pollen Robotics acquisition 10 None 2025-04-23
Parquet Content-Defined Chunking 10 None 2025-09-09
Phi-4 weights have been released under MIT license 10 None 2025-01-08
Show HN: A Transformer model that preserves logical equivalence 9 None 2025-03-02
MistralAI released a new Magistral Small 2509 8 None 2025-09-17
Z-Image Turbo Released – 6B Parameter Text to Image Model 8 None 2025-11-27
Tencent's Hunyuan Instruct 7B/4B/1.8B/0.5B new models have been released 8 None 2025-08-04
DeepSeek-Prover-V2-671B 8 None 2025-04-30
Sesame CSM-1B: Open-Source Conversational Speech Model 8 None 2025-03-14
Model Context Protocol (MCP) Course 8 None 2025-05-21
ByteDance/Dolphin on HuggingFace 7 None 2025-05-19
Hugging Face datasets and models for cybersecurity/sofwtare vulnerabilities 7 None 2025-03-09
Kimi-K2-Thinking: open weights LLM with frontier performance 7 None 2025-11-06
DeepSeek-v3.2-Exp 7 None 2025-09-29
Transformers v5 Is Out 7 None 2025-12-01
LFM2 WebGPU 7 None 2025-08-06
OpenAI/GPT-OSS-120B · Hugging Face 7 None 2025-08-05
Holo1.5: Foundational Models for Computer Use Agents 7 None 2025-09-15
Cybersecurity Instruction Tuned Model 6 None 2025-08-05
SigLIP 2: A better multilingual vision language encoder 6 None 2025-02-22
Qwen3-VL-30B-A3B-Instruct and Thinking 6 None 2025-10-04
New Open Source Technique shrinks LLMs to let them run on less … 6 None 2025-10-05
Granite 4.0 Nano: Just how small can you go? 6 None 2025-10-31
Weibo launch open source AI, VibeThinker-1.5B 6 None 2025-11-12
Better than DeepSeek R1? MiniMax-M1:open-weight hybrid-attention reasoning model 6 None 2025-06-16
More Efficient Chain-of-Thought Reasoning Through Certainty Probing 6 None 2025-02-18
Qwen2.5-Omni Technical Report 6 None 2025-03-30
DeepSeek-R1 without CCP censorship 6 None 2025-02-20
DeepSeek-v3.2 6 None 2025-09-29
Unlocking On-Policy Distillation for Any Model Family 6 None 2025-10-29
Apple releases FastVLM and MobileCLIP2 on HF, real-time video captioning 6 None 2025-08-30
Microsoft Phi 4 with R1 Reasoning 6 None 2025-02-04
Show HN: We built a better reranker and open sourced it 6 None 2025-08-27
Nvidia STT Parakeet v3 6 None 2025-08-15
First 70B model released with all training epochs and data 6 None 2025-09-12
Qwen3-Next series represents our next-generation foundation models 6 None 2025-09-12
Australian-made LLM beats OpenAI and Google at legal retrieval 6 None 2025-10-23
The 1B Token Challenge: Finding the Perfect Pre-Training Mix 6 None 2025-11-15
Kokoro-TTS 6 None 2025-01-13
Qwen Image Edit - SOTA Open Weight Image Editing Model 6 None 2025-08-18
Show HN: Agent Leaderboard 2.0 – Domain Specific edition 6 None 2025-07-17
Can you visualize what NYC smells like? Yes, turns out, you can 5 None 2025-12-18
Open R1: Update #2 5 None 2025-02-11
Deepseek VL2 Small 5 None 2025-02-08
Gemma 3 QAT (Quantized Aware Training) 3x less memory 5 None 2025-04-03
DocumentAI with 256M Parameters 5 None 2025-03-20
An open source common knowledge and context based Hallucination Detection Model 5 None 2025-04-29
Mixture of Tunable Experts-DeepSeek R1 Behavior Modification at Inference Time 5 None 2025-05-01
CircleGuardBench Leaderboard 5 None 2025-05-07
Show HN: Raman-01 – A Pocket Physics Solver LLM 5 None 2025-05-05
An MCP-powered agent in 50 lines of code 5 None 2025-05-15
SWE-rebench: Over 21,000 Open Tasks for SWE LLMs 5 None 2025-05-29
The Common Pile v0.1 5 None 2025-06-06
You could have designed state of the art positional encoding 5 None 2025-05-20
Show HN: KaniTTS – Open-source high-fidelity TTS with just 450M params 5 None 2025-09-19
GLM 4.5 5 None 2025-07-28
Gaia2 and Are: Empowering the Community to Evaluate Agents 5 None 2025-09-22
VibeVoice: A Frontier Open-Source Text-to-Speech Model 5 None 2025-08-26
Qwen2.5-Coder-3B Fine-Tuned for Triton Kernel Gen 5 None 2025-08-03
HuggingFace Skills: Fine-tune any LLM with one sentence for $0.30 5 None 2025-12-10
Show HN: Dante-Qwen-4B – Curing LLM "Neurosis" with a Divine Comedy Curriculum 5 None 2025-11-28
WindowSeat 5 None 2025-12-12
Microsoft open sources text-to-speech model VibeVoice‑Realtime‑0.5B 5 None 2025-12-04
DeepSeek-v3.2 and v3.2 Speciale Announced 5 None 2025-12-01
Ling-1T: 1T-parameter model with 50B active parameters per token 5 None 2025-10-08
Kimi K2: 1T total parameter open-source LLM by Moonshot AI 4 None 2025-07-11
Mistral AI releases Devstral-Small-2507 4 None 2025-07-10
Maya1: Open-source 3B Voice Model 4 None 2025-11-05
Apple: STARFlow-V, a Normalizing Flow Model for Causal Video Generation 4 None 2025-12-01
AI Energy Score v2: Refreshed Leaderboard, Now with Reasoning 4 None 2025-12-06
DeepMath: A lightweight math reasoning Agent with smolagents 4 None 2025-12-09
Qwen/QwQ-32B released on Hugging Face 4 None 2025-03-06
Wan2.1-T2V-14B 4 None 2025-02-25
The Curse of Depth in Large Language Models 4 None 2025-02-13
UIGEN-X-32B-0727 Reasoning Only UI Generation Model 4 None 2025-07-28
Pruned expert GPT-OSS 6.6B 4 None 2025-08-13
Gemma 3-270M 4 None 2025-08-14
OmniNeural – First NPU-Aware Multimodal Model 4 None 2025-08-24
Kimi-K2-Instruct-0905 4 None 2025-09-05
Tricks from OpenAI GPT-OSS you can use with transformers 4 None 2025-09-11
Show HN: Single-agent long-horizon reasoning within one LLM run 4 None 2025-07-23
Trackio: A new experiment tracking library from Hugging Face 4 None 2025-07-29
A 337M RSS feed dataset 4 None 2025-08-26
Migrating Hugging Face off Git LFS and to a new storage system … 4 None 2025-03-18
Care About Partial Differential Equations (PDEs) 4 None 2025-12-12
NitroGen: Unified vision-to-action model designed to play video games 4 None 2025-12-21
1M+ LoC, 400 models still afloat: The engineering behind Transformers 4 None 2025-10-06
AnyLanguageModel: One API for Local and Remote LLMs on Apple Platforms 4 None 2025-11-20
Show HN: Emoji Search – semantic emoji picker using sentence-transformers 4 None 2025-09-29
Granite-4.0-Micro: a 3.4B parameter LLM that runs in the browser 4 None 2025-10-06
Higgs – Rapidly Compress LLMs Without Significant Loss of Quality 4 None 2025-04-12
Open R1's OlympicCoder beats Deepseek R1, models and underlying dataset released 4 None 2025-03-25
Qwen2.5-Omni-7B 4 None 2025-03-26
MoCha: Towards Movie-Grade Talking Character Synthesis 4 None 2025-04-01
Show HN: HalluMix – A Benchmark for Real-World LLM Hallucination Detection 4 None 2025-05-06
New virtual try on model family that seems to be SOTA 4 None 2025-06-28
Gemma 3n available in the open-source ecosystem 4 None 2025-06-26
Automated Discovery of High-Performance GPU Kernels with OpenEvolve 4 None 2025-06-28
Jan-Nano-128k: Empowering deeper research through extended context understanding 4 None 2025-06-25
Paper2Agent: Research Papers as Interactive AI Agents 4 None 2025-10-10
Nvidia GPU Tflop Finder 4 None 2025-10-18
Mistral misspelled Ministral on HuggingFace and Ollama 4 None 2025-12-02
Maincoder-1B – an open 1B-parameter coding model with 76% HumanEval 4 None 2025-12-24
Ltxv-13B – high-quality videos in real-time 4 None 2025-05-07
Devin's First Open Source Model Beats O3 4 None 2025-05-06
Qwen 2.5 Max 4 None 2025-01-28
Hugging Face open sources a web-browsing agent that uses VLMs 4 None 2025-01-24
Deepseek R1 Zero 4 None 2025-01-20
LLaSE-G1 A FOSS speech enhancement model 4 None 2025-03-08
Kimi-Dev-72B 4 None 2025-07-13
WanX open weight sota 14B video model release 3 None 2025-02-25
Qwen3 235B (MoE with 128 experts) 3 None 2025-04-28
Black-forest-labs/FLUX.2-dev 3 None 2025-11-25
Collection of LLMs that run well in 32gb VRAM 3 None 2025-11-25
Epstein Files Dataset 3 None 2025-11-17
Xiaomi MiMo 3 None 2025-04-30
Yambda-5B – Industrial-scale music recommendation dataset 3 None 2025-06-04
Show HN: we released an open source, best-in-class medical reasoning model 3 None 2025-05-13
Understanding MCP Evals: Why Evals Matter for MCP 3 None 2025-06-06
Show HN: Ego-Dex Gradio App 3 None 2025-06-03
Hugging Face Courses 3 None 2025-05-27
Show HN: Tinker with Meta's "tokenizer-free" patcher 3 None 2025-05-21
Radiology explainer demo 3 None 2025-05-20
Memelang – a hybrid relational-graph query language 3 None 2025-05-17
GLM-4-32B-0414: New MIT-licensed SOTA LLM from Zhipu AI 3 None 2025-04-15
Drape1: Open-Source Scalable adapter for clothing generation 3 None 2025-05-01
Hugging Face Collaborates with Proxima Fusion on ML for Stellarator Optimization 3 None 2025-07-02
Largest in-person AV conversational dataset ever released 3 None 2025-06-27
Kimina-Prover: Applying Test-time RL Search on Large Formal Reasoning Models 3 None 2025-07-10
Show HN: 1.5B LLM routing model that aligns to preferences, not leaderboards 3 None 2025-07-17
Mistral Releases Voxtral: Open Source Speech Understanding Models (3B and 24B) 3 None 2025-07-15
CommaCarSegments: 3148 hours of raw CAN bus data from 230 different car … 3 None 2025-07-10
VACE: All-in-One Video Creation and Editing from Alibaba 3 None 2025-03-12
Nvidia Isaac GR00T N1 is the first open foundation model for humanoid 3 None 2025-03-21
DeepSeek V3-0324 Posted to HuggingFace 3 None 2025-03-24
AgentRxiv: Towards Collaborative Autonomous Research 3 None 2025-03-25
Training LLMs with GRPO and Interpreter Feedback Using WebAssembly 3 None 2025-04-06
EuroBERT: A High-Performance Multilingual Encoder Model 3 None 2025-03-10
Show HN: First large scale evaluation of 4o Image Generation from OpenAI 3 None 2025-03-27
Step-Audio-Chat: a 132B end-to-end speech-to-speech model 3 None 2025-02-17
Dia 1.6B – Nari Text-to-Speech Synthesis 3 None 2025-04-24
Microsoft Releases Phi-4-multimodal [pdf] 3 None 2025-02-26
Gelato-30B-A3B: A Grounding Model for GUI Computer-Use Tasks 3 None 2025-11-11
Ring-1T: open-source, SOTA thinking model with a trillion parameters 3 None 2025-11-11
New Nvidia Nemotron models – new king of local models? 3 None 2025-10-31
Supercharge Your OCR Pipelines with Open Models 3 None 2025-10-31
MTEB v2: Evaluation of embedding and retrieval systems for more than just … 3 None 2025-10-20
AnyCoder creates a demo for Qwen Image Edit Plus in 10mins 3 None 2025-09-22
I made WEBGEN-OSS-20B, a model that generates clean websites from your prompts 3 None 2025-09-13
Reasoning Traces from QA Pairs 3 None 2025-09-09
Welcome EmbeddingGemma, Google's new efficient embedding model 3 None 2025-09-04
Output Schema for CodeAct AI Agents: From Trial-and-Error to Predictive Planning 3 None 2025-08-31
WildChat-4.8M: 4.8M Real User–ChatGPT Conversations (Open Dataset) 3 None 2025-08-11
Break the quadratic wall of Transformer attention: WERSA, paper+code open source 3 None 2025-08-02
Qwen-Image-Edit-2509 3 None 2025-09-22
AI Spreadsheet Benchmark [pdf] 3 None 2025-09-22
FinePDFs Dataset 3 None 2025-09-15
TildeOpen-30B: European LLM Focused on Underrepresented Languages 3 None 2025-09-04
First vision language model built off Open AI GPT-OSS 3 None 2025-08-26
Seed-OSS: open-source LLM models by ByteDance 3 None 2025-08-22
From Zero to GPU: A Guide to Building and Scaling Production-Ready CUDA … 3 None 2025-08-20
Jan-v1: Advanced Agentic Language Model 3 None 2025-08-12
NextCoder by Microsoft — LLM performing on par with GPT-4o on complex … 3 None 2025-08-08
OpenReasoning-Nemotron by Nvidia: state-of-the-art distilled reasoning models 3 None 2025-08-08
Accelerate ND-Parallel: A Guide to Efficient Multi-GPU Training 3 None 2025-08-08
GEN3C: 3D-Informed World-Consistent Video 3 None 2025-03-06
DeepSeek-R1 on iPhone? (DeepSeek-R1-Distill-Qwen-1.5B-GGUF) 3 None 2025-01-21
HuggingFace open reproduction of R1 data and training pipeline 3 None 2025-01-27
Hugging Face AI Agents Course 3 None 2025-02-10
Fine-Tune Deepseek-R1 with a Synthetic Reasoning Dataset 3 None 2025-02-11
Jamba Reasoning 3B 3 None 2025-10-11
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention 3 None 2025-10-09
The First Long Context Guardrail 3 None 2025-10-09
Apriel-1.5-15B-Thinker 3 None 2025-10-01
In Defense of Tokenizers 3 None 2025-09-28
Timeline of AI model releases in 2024 3 None 2025-01-01
SAM3D-Body with glb export in rerun 3 None 2025-11-28
Building Fast Vector Search for Legal Documents 3 None 2025-10-20
Ant releases the first open-source trillion-parameter inference model, Ring-1T 3 None 2025-09-30
Speech to text model for healthcare-based voice applications 3 None 2025-12-21
The Jagged AI Frontier Is a Data Frontier 3 None 2025-12-17
Show HN: Reasoning models don't guarantee better security 3 None 2025-12-16
Enhancing LLMs with LoRA – Standardized Recipes for Capability Enhancement 3 None 2025-12-16
Circuit Sparsity 3 None 2025-12-13
AutoGLM-Phone-9B-Multilingual: Vision-language model for automated mobile agents 3 None 2025-12-11
Transformers Are Multi-State RNNs 3 None 2025-12-08
The LLM Evaluation Guidebook 3 None 2025-12-04
Microsoft/MAI-DS-R1, DeepSeek R1 Post-Trained by Microsoft 3 None 2025-04-18
N-Atlas V1 2 None 2025-09-21
Statistical Methods in Generative AI 2 None 2025-09-16
EmbeddingGemma is a 300M parameter, open embedding model from Google 2 None 2025-09-05
Swiss AI Initiative 2 None 2025-09-02
Apertus LLM 2 None 2025-09-02
Hugging Face speadsheet tool: AI Sheets 2 None 2025-09-01
A Novel Pretrained Tokenizer-Free LLM Architecture 2 None 2025-08-29
MiniCPM-V 4.5: GPT-4o Level MLLM for Image and Video Understanding on Your … 2 None 2025-08-26
NASA and IBM release open source model on Hugging Face to predict … 2 None 2025-08-20
Tokenizers 2 None 2025-08-17
FormulaOne: A reasoning benchmark that all models score 0% on 2 None 2025-08-14
dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model 2 None 2025-08-06
Qwen3-30B-A3B-Thinking-2507 has been released 2 None 2025-07-31
Intern-S1: A 241B parameter open-source MoE multimodal model 2 None 2025-07-28
Creating custom kernels for the AMD MI300 2 None 2025-07-25
Fast LoRA Inference for Flux with Diffusers and PEFT 2 None 2025-07-24
Nvidia parakeet-tdt-0.6B-v2 2 None 2025-07-22
How to Run a Hugging Face Model in Jax (Part 1) 2 None 2025-07-20
Show HN: Chimera-QxD-BMM-Qwen2-l22_28-alphaqd-1.5B-f16 2 None 2025-07-19
Phi-4-Reasoning 2 None 2025-05-01
Building the Hugging Face MCP Server 2 None 2025-07-10
Kimi-K2-Base 2 None 2025-07-11
Metalorian: Generate Heavy Metal-Binding Peptides with Diffusion Sampling 2 None 2025-07-12
Gemma3 on Hugging Face 2 None 2025-03-26
Show HN: KaniTTS – Ultra Fast and Expressive TTS Model 2 None 2025-09-22
Granite docling 258M: a small multimodal model for efficient document conversion 2 None 2025-09-17
DeepSeek-R1 WebGPU 2 None 2025-01-22
Bespoke-Stratos-17k: Open Reasoning Dataset by Distilling DeepSeek-R1 2 None 2025-01-27
The state of open video generation models 2 None 2025-01-28
Generate Images, Chat with PDF in WebGPU via DeepSeek Janus Pro 1B 2 None 2025-01-28
Mistral-Small-24B-Base-2501 2 None 2025-01-30
#9: Does AI Remember? The Role of Memory in Agentic Workflows 2 None 2025-02-03
FinePersonas 2 None 2025-02-10
OpenAI o3 just scored 99.8% on CodeForces using brute-force 2 None 2025-02-12
FantasyTalking: Realistic Talking Portrait Generation 2 None 2025-04-30
Neural Network Visualizer 2 None 2025-04-29
The Bitter Lesson Learned from 2k Multilingual Benchmarks 2 None 2025-04-23
ThinkFlow: The Revolutionary Platform That Gives LLMs the Power to Think 2 None 2025-04-19
Open-source LLM beats OpenAI o1 and DeepSeek-R1 for PyTorch-to-Triton codegen 2 None 2025-03-19
Cohere: Command A (111B Open Weights Model) 2 None 2025-03-14
Open Dataset: Vehicle Accidents 2 None 2025-03-13
Flex.1-Alpha – A new modded Flux model that can properly handle being … 2 None 2025-01-19
ModernBERT: Encoder-only Transformer Model Strictly Improving on past work 2 None 2025-01-01
Hugging Face advocates for Code Agents: agents that write tool calls as … 2 None 2025-01-02
JFK Assassination Records Dataset on Hugging Face 2 None 2025-04-09
Forget What You Know about LLMs Evaluations – LLMs Are Like a … 2 None 2025-02-13
Desklib AI Detector Ranks No 1 on Raid Benchmark for AI Detection 2 None 2025-02-17
SWE-Lancer: Can LLMs Earn $1M from Real-World Freelance Software Engineering? 2 None 2025-02-18
Show HN: Roast Any Website with AI 2 None 2025-02-25
FastRTC: The Real-Time Communication Library for Python 2 None 2025-02-25
Hugging Face Smolagents 2 None 2025-01-05
Show HN: We collected detailed annotations for text-to-image generation 2 None 2025-01-10
Vdr-2B-multi-v1 a multilingual embedding model for visual document retrieval 2 None 2025-01-10
Microsoft BitNet 1.58bit LLM 2B4T released 2 None 2025-04-16
MamayLM: An Efficient Ukrainian LLM 2 None 2025-04-23
HuggingFace on Sheets 2 None 2025-03-24
WebThinker: Empowering Large Reasoning Models with Deep Research Capability 2 None 2025-05-01
Show HN: TTS Arena V2 2 None 2025-05-02
Show HN: My progress towards building a robotics training dataset 2 None 2025-03-18
HOGWILD! Inference – parallel LLM chain-of-thought with shared attention 2 None 2025-04-09
Llama-4 Model-Based Agentic AI System HuggingFace Released 2 None 2025-04-06
Llama 3.2 from-scratch implementation focused on code readability 2 None 2025-04-01
deepsite 2 None 2025-03-31
SuperBPE: Space Travel for Language Models 2 None 2025-03-29
HuggingChat is shutting down (for now) 2 None 2025-07-04
Skywork-R1V3-38B open-source multimodal reasoning model 2 None 2025-07-08
A Survey on Latent Reasoning 2 None 2025-07-10
Show HN: AEE – An Open-Source Engine That Evaluates Truth and Bias … 2 None 2025-04-13
Magi-1: Autoregressive Video Generation at Scale 2 None 2025-05-06
Veena – open-source TTS for Indian Languages 2 None 2025-06-25
FLUX Kontext Dev Ultra Fast Live 2 None 2025-06-26
Building and better understanding vision-language models (2024) 2 None 2025-05-10
Vision Language Models (Better, Faster, Stronger) 2 None 2025-05-13
Show HN: 2.4x faster baai/bge-M3 2 None 2025-05-18
Tiny Agents in Python: an MCP-powered agent in ~70 lines of code 2 None 2025-05-23
The Qwen3 Embedding Model 2 None 2025-06-06
MiniCPM4 – a series of open multimodal models for edge inference 2 None 2025-06-10
Embedding Benchmark for Retrieval 2 None 2025-06-11
Wan: Open and Advanced Large-Scale Video Generative Models 2 None 2025-05-14
KernelLLM – Meta's new 8B SotA model 2 None 2025-05-19
How do AI political biases differ between English and French? 2 None 2025-05-21
TiRex Leads Gift Eval 2 None 2025-06-02
SOTA Model in 8B Size? 2 None 2025-05-29
The 4 Things the Qwen-3's Chat Template Teaches Us 2 None 2025-05-02
Show HN: A synthetic text dataset to train tiny language models on 2 None 2025-05-01
Qwen3Guard: Real-Time Safety for Your Token Stream 2 None 2025-09-24
K2-Think: A Parameter-Efficient Reasoning System 2 None 2025-09-13
Environments Hub: Your Language Model needs better (open) environments to learn 2 None 2025-09-05
Pivotal Token Search (PTS): Targeting Critical Decision Points in LLM Training 2 None 2025-08-18
Voxtral WebGPU 2 None 2025-07-25
Show HN: kulyk-uk-en and kulyk-en-uk 2 None 2025-07-22
FP8 DeepSeek R1 Distilled LLMs for SGLang and VLLM 1 None 2025-01-29
Show HN: An Agentic AI dataset for deepfake detection 1 None 2025-01-15