August 2026 Summaries
4 posts from Momento
Filter
Month:
Year:
Post Summaries
Back to Blog
Benchmarking Valkey Search showed that the number of indexes alone is not the main determinant of write performance: indexes on unrelated key prefixes impose no measurable cost because prefix matching uses a trie, while multiple indexes matching the same keys create separate indexing jobs and reduce write capacity, though with diminishing marginal impact as work is distributed across writer threads. In tests on an r7i.4xlarge server, adding a first standard TAG and NUMERIC index reduced sustainable writes from 80,000 to about 33,636 per second, while four matching indexes sustained about 16,818 writes per second but processed over 67,000 indexing jobs per second. Vector indexing was far more costly, as adding a 1024-dimensional HNSW field dropped throughput from 32,000 to 2,828 writes per second and produced sharp latency cliffs once capacity was exceeded. However, adding a lightweight lexical exact-match index alongside a vector-heavy semantic index caused only about a 3% additional impact because writes wait primarily for the slower vector operation. The findings suggest capacity planning should focus on the fields in indexes that match hot keys, especially vector fields, rather than on overall index count, and emphasize measuring queue depth and latency at progressively varied load levels to identify sustainable throughput before saturation.
Aug 27, 2026
1,863 words in the original blog post.
Agent-memory retrieval with Valkey Search requires balancing vector similarity with metadata such as recency and task outcomes, since filters influence not only which memories qualify but also how searches are executed. Valkey’s query planner chooses between an exact pre-filtered scan for filters estimated to match 0.1% or fewer indexed vectors and inline filtering during HNSW traversal for broader matches, making index size, tenant scoping, and filter selectivity important performance considerations. Because the threshold is relatively strict and not normally configurable, the author recommends structuring indexes and key prefixes to isolate tenants or other logical namespaces, reducing qualifying sets and enabling fast exact searches. The discussion also highlights that deleted, expired, or evicted vectors remain as stranded nodes in the HNSW graph until the index is rebuilt, while in-place vector updates are comparatively inexpensive. To avoid graph bloat caused by creating TTL-based keys for every completed task, agent memories should use stable task-derived keys and be updated over time, with query-time timestamp ranges serving as the main recency mechanism and TTLs used more conservatively for storage control.
Aug 19, 2026
1,559 words in the original blog post.
Momento has opened a limited preview of Cluster and Flex configurations for Momento Cache, offering fully managed, single-tenant Valkey capacity for teams needing greater isolation, control, and performance than its Serverless option. Cluster lets users specify instance types, shard counts, replicas, and availability zones, while Flex automatically optimizes resources within user-defined limits; Momento manages provisioning, failover, scaling, upgrades, patches, and topology changes. The service uses a gateway-based architecture that provides a stable RESP endpoint compatible with Redis and Valkey clients while handling traffic challenges such as hot keys, connection storms, TLS, authentication, rate limits, and request coalescing before requests reach Valkey nodes. Pricing is based on deployed capacity without separate request, transfer, or connection charges, with Flex priced per GiB-month and Cluster priced per deployed instance, and the platform supports autoscaling and optional bring-your-own-cloud deployments. Users can request preview access, configure the CLI, create a capacity pool and database, and connect through TLS, while future plans include fine-grained access controls, VPC peering, and S3 integration.
Aug 18, 2026
1,021 words in the original blog post.
Buffer-Free Video AI’s second edition convened 117 of 124 registered attendees from 41 organizations, achieving a 94.4% attendance rate that organizers viewed as evidence of strong engagement among streaming and media professionals. The two-day event emphasized practitioner-led, non-promotional technical discussion, with 22 sessions featuring 29 speakers from 17 companies and 40% of speakers coming from non-vendor organizations including Paramount, Netflix, FOX, CBS Interactive, Meta, and Red Bull Media House. Presentations focused on scaling production and live-video infrastructure, practical applications and limitations of AI in media workflows, and the operational challenges of playback across fragmented devices and platforms. Examples included FIFA World Cup-scale streaming planning, FAST-channel expansion, Netflix’s high-volume infrastructure, AI-powered sports content discovery, accelerated encoding, automated quality control, and anti-piracy tools. Interactive workshops, networking, and detailed audience Q&A were presented as central to the event’s value, enabling participants to exchange candid lessons from real production systems.
Aug 05, 2026
1,201 words in the original blog post.