Home / Companies / Modal / Blog / March 2024

March 2024 Summaries

2 posts from Modal

Filter
Month: Year:
Post Summaries Back to Blog
Ramp, a spend management company, was struggling with fine-tuning its large language models (LLMs) and scaling batch processing. They initially tried using LLM providers like OpenAI but were limited by customizability concerns and high costs. Ramp then adopted Modal, a platform that allowed them to fine-tune their models while controlling each step of the fine-tuning workflow. By using Modal, Ramp was able to accelerate development of their text-to-structured-JSON model for receipt management, driving down receipts requiring manual intervention by 34%. Modal's serverless platform also enabled Ramp to parallelize tasks and speed up LLM batch processing, resulting in significant cost savings and productivity gains. With Modal as part of its data processing stack, Ramp is now able to ship its AI features faster than ever before.
Mar 26, 2024 517 words in the original blog post.
Modal developed modal-http, a Rust-based HTTP and WebSocket service that translates incoming web traffic into its existing serverless function-call system, allowing users to deploy endpoints ranging from simple APIs to long-running, GPU-intensive workloads with large inputs and outputs. Unlike many serverless platforms with strict execution, memory, bandwidth, and payload limits, Modal supports containers with up to 64 CPUs, 336 GB of memory, and eight Nvidia H100 GPUs, requiring an architecture that can efficiently autoscale while streaming large requests and responses. The service represents HTTP interactions as serialized ASGI-style events, enabling request and response streaming, backpressure, error propagation, request bodies up to 4 GiB, and unlimited streamed responses. It handles browser idle timeouts for lengthy work through periodic HTTP redirects and extends the same event-based model to WebSocket handshakes and bidirectional messaging. Modal runs modal-http behind a TCP load balancer, Caddy, and Kubernetes for conventional ingress infrastructure, while its custom serverless runtime handles compute separately; it currently serves ingress from Ashburn, Virginia, with additional regions planned. The company reports that replacing its prior Python-based ingress with this Rust implementation reduced production 502 errors by 99.7% and established shared infrastructure for web functions and remote function calls.
Mar 14, 2024 3,883 words in the original blog post.