Home / Companies / Supermemory / Blog / Post Details
Content Deep Dive

Building code-chunk: AST Aware Code Chunking

Blog post from Supermemory

Post Details
Company
Date Published
Author
Shoubhit
Word Count
2,007
Company Posts That Month
2
Language
English
Hacker News Points
2
Post removed?
No
Summary

Supermemory introduces code-chunk, an AST-based library for preparing source code for retrieval-augmented generation systems, arguing that character-based text splitting often breaks functions and loses meaningful programming context. Built on the cAST research approach and tree-sitter parsers, the library identifies semantic entities such as functions, classes, methods, imports, signatures, documentation, and parent-child scope relationships, then recursively creates size-bounded chunks at syntactic boundaries using non-whitespace character counts. It enriches chunks with contextual metadata including file paths, scopes, defined signatures, dependencies, and nearby code, intended to improve embedding quality and retrieval relevance. Beyond the research prototype, code-chunk provides overlap options, streaming and batch processing, Effect integration, and WASM support for edge environments, while supporting TypeScript, JavaScript, Python, Rust, Go, and Java. The authors also revised their evaluation setup with same-repository distractor files and intersection-over-union relevance thresholds, reporting stronger Recall@5 and IoU@5 results than fixed-size and alternative code chunkers, as well as reduced time, tokens, cost, and tool calls in a SWE-bench Lite agent experiment using semantic search.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Vector Search 12 1,607 321 133 +4%
RAG 5 974 222 101 -17%
Real-time 1 8,461 1,407 260 +57%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.