Home / Companies / Together AI / Blog / November 2023

November 2023 Summaries

5 posts from Together AI

Filter
Month: Year:
Post Summaries Back to Blog
Together AI has raised $102.5M in a Series A financing to build a comprehensive cloud platform for open source AI models, with lead investors including Kleiner Perkins, NVIDIA, and Emergence Capital. The new capital will accelerate the development of the fastest cloud platform for generative AI applications, allowing developers to easily integrate leading open source models or create their own models through pre-training or fine-tuning. With a focus on core research, Together AI publishes novel research under open-source licenses that benefit the AI ecosystem, and has released several notable datasets and products, including RedPajama-V2, which is widely adopted by AI researchers and developers. The company aims to create a platform technology that provides choice and options for the AI ecosystem, enabling any researcher or developer to participate in shaping its future.
Nov 29, 2023 895 words in the original blog post.
FlashFFTConv is an algorithm for efficiently computing FFT convolutions on GPUs, which speeds up convolutions by up to 7.93x over PyTorch and achieves up to 4.4x speedup end-to-end. It addresses the bottlenecks of traditional FFT algorithms in machine learning hardware, particularly I/O and matrix-matrix multiply operations. The algorithm uses a Monarch decomposition of the FFT, which breaks down the convolution into matrix-matrix multiply operations that can be efficiently computed on tensor cores. This allows for faster convolutions and better scaling with sequence length, making it suitable for long-sequence tasks such as audio analysis and DNA modeling. FlashFFTConv has been integrated into research codebases and is expected to enable new applications in machine learning.
Nov 13, 2023 1,804 words in the original blog post.
Together Custom Models is an end-to-end solution for building powerful LLMs from data design to evaluation, providing a state-of-the-art cluster with NVIDIA H100 and A100 GPUs, and a team of expert researchers available to work with customers every step of the way. The models built with Together Custom Models are owned by the customer, allowing for full ownership, control, and customization. The solution includes stages such as data discovery & optimization, selecting model architecture, hyperparameters, and training strategy, training, tuning & alignment, and evaluation. It leverages state-of-the-art techniques like DSIR, DoReMi, FlashAttention-2, and CocktailSGD to achieve fast and reliable performance. With Together Custom Models, customers can build custom models that are tailored to their specific needs, resulting in higher accuracy and adaptability to their tasks, with the best price-performance tradeoff for production applications.
Nov 13, 2023 1,289 words in the original blog post.
The Together Inference Engine is a fast inference stack that outperforms other services by up to 3x when running on the same hardware, with performance of 117 tokens per second on Llama-2-70B-Chat and 171 tokens per second on Llama-2-13B-Chat. It is built on CUDA and runs on NVIDIA Tensor Core GPUs, utilizing techniques such as FlashAttention-2, Flash-Decoding, and Medusa to optimize inference performance. The engine achieves results comparable to the reference Hugging Face implementation without compromising quality. New features include Serverless Endpoints with automatically added capacity and scaling, Dedicated Instances for custom models, Auto-scaling for increased flexibility, and expanded model availability. Pricing has been lowered due to efficiency gains, making it 20% cheaper than some competitors while offering faster performance.
Nov 13, 2023 880 words in the original blog post.
Together GPU Clusters offers purpose-built dedicated GPU training clusters for startups and enterprises to accelerate generative AI development, providing cutting-edge hardware, an optimized software stack, flexible capacity, and top-tier support. The clusters deliver unparalleled model training speeds, amazing cost efficiency, and expert support, meeting the key needs of customers who require high-performance computing infrastructure to train gen AI models. With its state-of-the-art NVIDIA GPUs, fast Infiniband networking, and optimized software stack, Together GPU Clusters helps customers like Pika Labs and NexusFlow scale their generative AI development and achieve significant cost savings and time reductions.
Nov 13, 2023 929 words in the original blog post.