March 2024 Summaries
2 posts from Deepinfra
Filter
Month:
Year:
Post Summaries
Back to Blog
DeepInfra has introduced "JSON mode" across its text language models, allowing for outputs that conform to valid JSON without any performance overhead. This feature is activated by setting the "response_format" parameter to {"type": "json_object"} in various text APIs and is available for free on all models. JSON mode is designed to produce concise, data-driven responses, skipping unnecessary text, and is particularly suited for structured data requests such as dates or lists. However, it may result in the models generating information if they lack the necessary details to answer a query accurately. This enhancement is part of DeepInfra's efforts to improve language model functionality, and feedback is encouraged to further refine the capability.
Mar 08, 2024
624 words in the original blog post.
DeepInfra offers a platform for deploying custom language models with a simple API and predictable pricing, allowing users to host models on Hugging Face with options for private repositories to enhance security. Users can deploy models via a web interface or an HTTP API, specifying various settings such as GPU type, number of GPUs, and model repository details. DeepInfra provides fully managed GPU infrastructure for running models at scale, promising enterprise-grade uptime at competitive rates, and features various models, including popular ones like OpenAI's GPT and Meta's LLaMA. Comprehensive documentation and sales support are available for users seeking to leverage these AI hosting solutions.
Mar 01, 2024
276 words in the original blog post.