Building code-chunk: AST Aware Code Chunking
Blog post from Supermemory
Supermemory introduces code-chunk, an AST-based library for preparing source code for retrieval-augmented generation systems, arguing that character-based text splitting often breaks functions and loses meaningful programming context. Built on the cAST research approach and tree-sitter parsers, the library identifies semantic entities such as functions, classes, methods, imports, signatures, documentation, and parent-child scope relationships, then recursively creates size-bounded chunks at syntactic boundaries using non-whitespace character counts. It enriches chunks with contextual metadata including file paths, scopes, defined signatures, dependencies, and nearby code, intended to improve embedding quality and retrieval relevance. Beyond the research prototype, code-chunk provides overlap options, streaming and batch processing, Effect integration, and WASM support for edge environments, while supporting TypeScript, JavaScript, Python, Rust, Go, and Java. The authors also revised their evaluation setup with same-repository distractor files and intersection-over-union relevance thresholds, reporting stronger Recall@5 and IoU@5 results than fixed-size and alternative code chunkers, as well as reduced time, tokens, cost, and tool calls in a SWE-bench Lite agent experiment using semantic search.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Vector Search | 12 | 1,607 | 321 | 133 | +4% |
| RAG | 5 | 974 | 222 | 101 | -17% |
| Real-time | 1 | 8,461 | 1,407 | 260 | +57% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.