April 2024 Summaries
3 posts from Together AI
Filter
Month:
Year:
Post Summaries
Back to Blog
Together AI has partnered with Snowflake to bring Arctic LLM, an enterprise-grade large language model, to Enterprise customers through its Together Inference platform. The LLM is designed to be the most open and scalable on the market, featuring a unique Mixture-of-Experts architecture that delivers top-tier intelligence with unparalleled efficiency at scale. With this release, enterprises can now build and deploy production applications with complex workloads in their chosen environment, including private cloud or on-premise deployments. The model is optimized for industry benchmarks across various use cases and features a 480 billion parameter MoE model. Open-source models like Snowflake Arctic offer advantages to enterprises, including high accuracy, faster performance, privacy, flexibility of deployment configurations, and greater control over the model's operation.
Apr 25, 2024
422 words in the original blog post.
The Together AI Python SDK has been officially released with version v1, providing improved OpenAI compatible APIs for inference, fine-tuning, and other applications. The new API is more intuitive, supports async operations, and includes better error handling. Users can now stream responses from chat models, run completions on code and language models, use image models, generate embeddings, and fine-tune models with their own data. Fine-tuning is supported through the SDK or CLI, including the Llama 3 models, and users are encouraged to filter out low-quality data when using the RedPajama-V2 Dataset. The Python library is available on GitHub, and a similar TypeScript SDK is expected to be released in the coming weeks.
Apr 22, 2024
361 words in the original blog post.
Together AI has partnered with Meta to release Meta Llama 3, an accessible large language model designed for developers, researchers, and businesses to build generative AI applications. The model offers significant advancements in performance and capabilities, including improved reasoning, code generation, and following instructions. The Together Inference Engine provides industry-leading performance up to 350 tokens per second, enabling enterprises to build production applications in their chosen environment. The release also features pretrained and instruction fine-tuned language models with various parameter counts, as well as a safety model called LlamaGuard-V2-8B. The model's open-source nature is designed to promote responsible innovation and encourages users to experiment and scale their generative AI ideas responsibly.
Apr 18, 2024
602 words in the original blog post.