Home / Companies / Speedscale / Blog / March 2025

March 2025 Summaries

3 posts from Speedscale

Filter
Month: Year:
Post Summaries Back to Blog
LLM application testing should extend beyond model capability benchmarks and training concerns to emphasize production performance, including latency, throughput, error rates, saturation, rate limits, and cost. Using an example application with OpenAI chat completion and image-generation features, the discussion shows that external model calls can account for most request time, with chat responses around 1.5 seconds and image generation near 10 seconds, making endpoint context and asynchronous page design important. It recommends capturing real API traffic, then replaying it through service mocks to create repeatable tests that isolate application behavior from nondeterministic LLM responses, provider rate limits, outages, and per-request charges. Load tests can then increase virtual users, adjust infrastructure settings such as replicas and CPU or memory allocation, and inject failures to identify scaling limits and resilience weaknesses. Results should be evaluated through SRE-style golden signals and response-time percentiles, while teams balance model quality against speed and operational cost when selecting models for production.
Mar 21, 2025 2,602 words in the original blog post.
Five years of production gRPC experience indicate that its performance and streaming capabilities come with operational challenges rooted largely in HTTP/2, binary Protocol Buffers, and asynchronous connection behavior. End-to-end TLS requirements, inconsistent H2C upgrades, status codes stored in often-hidden HTTP/2 trailers, and incomplete proxy or firewall support can complicate routing and troubleshooting. Protobuf payloads improve efficiency but require schema files and specialized tools such as grpcurl, Wireshark, or protocol-aware proxies for inspection, while gRPC-Web may require gateways or careful feature-compatibility planning for browser clients. Converting protobuf messages to JSON can also cause confusion because default-valued fields are commonly omitted. Connection management requires particular attention, as long-lived streams may hang, intermediaries can unexpectedly reset traffic, and a successful SendMsg call confirms queuing rather than actual delivery. Effective production use therefore depends on HTTP/2-aware observability, tracing, documented timeout and retry configurations, standardized debugging tools, and realistic testing of outages, restarts, and reconnection behavior.
Mar 19, 2025 1,619 words in the original blog post.
Traditional Test Data Management relies on periodically copying and masking production databases for test environments, but the approach can create stale data, high storage and maintenance costs, security risks, and difficulties supporting rapidly changing cloud-native and microservices architectures. Production Traffic Replay is presented as a streaming alternative that captures live production requests, applies Data Loss Prevention rules before data reaches testing systems, and replays sanitized request sequences to provide current and behaviorally realistic test scenarios. By operating at the network or protocol level rather than relying on complete database copies, PTR can reduce infrastructure duplication, automate data protection and updates, support varied back-end technologies, and simplify performance, regression, and integration testing. The article argues that organizations can adopt PTR incrementally, beginning with individual services or endpoints, to complement or replace parts of existing TDM processes as release cycles and distributed systems become more complex.
Mar 12, 2025 1,345 words in the original blog post.