Home / Companies / Deepinfra / Blog / March 2023

March 2023 Summaries

3 posts from Deepinfra

Filter
Month: Year:
Post Summaries Back to Blog
Flan-UL2, a large open-source chatbot model with 20 billion parameters, offers a powerful alternative to proprietary AI models, providing users with the ability to deploy it easily via DeepInfra’s platform. This fine-tuned version of the UL2 model utilizes the Flan dataset and, due to its size, is challenging to run on personal hardware. Instead, it can be deployed on DeepInfra's managed infrastructure, where users only pay for inference time at a rate of $0.0005 per second, amounting to approximately $0.0001 per token generated. The deployment process is simplified through DeepInfra's web dashboard or API, allowing users to initiate inferences without the need for complex setups involving Docker or machine learning frameworks. This setup ensures a cost-effective and user-friendly approach to running advanced AI models, supported by DeepInfra’s reliable GPU infrastructure and customer assistance via Discord.
Mar 17, 2023 495 words in the original blog post.
DeepInfra provides a platform for running text-to-image models like runwayml/stable-diffusion-v1-5, allowing users to generate images via an API by submitting prompts. The service offers advanced options and settings, accessible through their model page or API documentation, and is designed to support scalable model execution with managed GPU infrastructure, emphasizing high uptime and cost-effectiveness. DeepInfra also features various models, including those from notable developers like OpenAI and Google, and provides resources such as pricing, documentation, and support to facilitate user engagement with their AI hosting services.
Mar 08, 2023 218 words in the original blog post.
DeepInfra offers a platform that allows users to deploy and run AI models at scale using a fully managed GPU infrastructure, providing enterprise-grade uptime at competitive rates. To utilize DeepInfra's services, users need to obtain an API key, which is necessary for all API requests and can be generated by signing up and accessing the dashboard. Model deployment can be done through the web dashboard or via API, and models are automatically deployed upon the first inference request. Users can perform model inference using DeepInfra's REST API, exemplified by using curl commands, to interact with various models like OpenAI's Whisper and others. DeepInfra also highlights its range of available models and provides additional resources such as pricing, documentation, and customer support to enhance user experience.
Mar 02, 2023 278 words in the original blog post.