Light
Home
/
Companies
/
Together AI
/
Hacker News
Together AI on HN
49 posts with 1+ points since 2022
Filters
Min points:
1
10
25
50
100
250
500
Since:
2023
2024
2025
2026
Posts by Month (49 total)
Hacker News Posts
Search:
Title
Points
Comments
Date
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-Precision
287
60
2024-07-11
RedPajama v2 Open Dataset with 30T Tokens for Training LLMs
236
60
2023-10-30
Paving the way to efficient architectures: StripedHyena-7B
221
65
2023-12-08
Consistency diffusion language models: Up to 14x faster, no quality loss
219
96
2026-02-20
AdapTive-LeArning Speculator System (ATLAS): Faster LLM inference
198
47
2025-10-12
Based: Simple linear attention language models
165
14
2024-03-05
Dragonfly: A large vision-language model with multi-resolution zoom
143
36
2024-06-06
Llama 32K Context Released by Together AI
84
9
2023-07-29
Together AI raises a $102.5M Series A
70
24
2023-11-29
Llama 2 on togetherAI is as bad of a privacy nightmare as …
54
22
2023-09-08
Mamba-3
50
1
2026-03-18
Direct Preference Optimization vs. RLHF
37
1
2025-05-25
DeepCoder: An Open-Source 14B Coder at O3-Mini Level
31
4
2025-04-09
The Mamba in the Llama: Distilling and Accelerating Hybrid Models
4
0
2024-09-09
Together Inference Engine 2.0 with new Turbo and Lite endpoints
3
0
2024-07-18
Fine-tuning Llama-3 to get 90% of GPT-4's performance at a fraction of …
3
0
2024-07-19
Fine-Tuning LLMs for Multi-Turn Conversations: A Technical Deep Dive
3
0
2024-11-27
Together AI acquires CodeSandbox to launch code interpreter for generative AI
3
0
2024-12-12
Together API hosts open source models
2
0
2023-07-14
Together Inference Engine – the fastest inference available
2
0
2023-12-12
Weak models excel at long context tasks
2
0
2026-03-27
Generate react apps with Llama 3.1
2
1
2024-08-02
Cursor and Together AI deliver real-time, low-latency inference at scale
2
0
2026-07-28
A/B testing LLMs in production
2
0
2026-08-18
Evo: Long-context modeling from molecular to genome scale
2
0
2024-02-27
AI for Systems: Using LLMs to Optimize Database Query Execution
2
0
2026-04-11
Together MoA–collective intelligence of open-source models pushing LLM frontier
2
0
2024-06-15
Instrumentation checklist for running large GPU clusters
2
2
2024-08-14
Speculative decoding for high-throughput long-context inference
2
0
2024-09-05
The Open Source AI Stack
2
0
2026-09-10
Parcae: Doing more with fewer parameters using stable looped models
2
0
2026-04-17
Llama-2-7B-32K-Instruct – and fine-tuning for Llama-2 models with Together API
2
0
2023-08-22
FlashAttention-2: Faster attention with better parallelism and work partitioning
2
0
2023-07-17
Free Llama 3.2 vision API
1
0
2024-09-25
New SOTA Reranker from Salesforce
1
0
2024-09-10
A universal subject Llama 3.1 tutor from Together AI
1
1
2024-08-03
Flux API available on Together AI:FLUX1.1 [pro] and free access FLUX.1 [schnell]
1
1
2024-10-03
Inference Optimization for MiniMax Sparse Attention
1
0
2026-06-03
Linearizing LLMs with LoLCATs
1
0
2024-10-15
Together Code Sandbox
1
0
2025-05-20
Together AI embeddings endpoint with higher quality, 4x lower cost than OpenAI
1
1
2024-01-11
The Frontier Is Open
1
0
2025-06-09
Fine-tuning open LLM judges to outperform GPT-5.2
1
0
2026-02-04
Cache-aware prefill–decode disaggregation – 40% faster long-context LLM serving
1
0
2026-02-12
Flash Attention 4
1
0
2026-03-05
CoderForge-Preview: SOTA open dataset for training efficient coding agents
1
0
2026-02-25
DSGym: A holistic framework for evaluating and training data science agents
1
0
2026-02-25
Aurora
1
0
2026-04-17
RedPajama-Data-v2: An open dataset with 30T tokens (2023)
1
0
2024-04-22