August 2026 Summaries
9 posts from Semgrep
Filter
Month:
Year:
Post Summaries
Back to Blog
Dependency cooldowns restrict package managers from installing versions released within a defined period, reducing exposure to newly published malicious dependencies while typically causing little developer disruption. The author describes implementing one-week cooldowns at Semgrep using Python’s uv package manager, where configuration settings require a sufficiently recent uv version and allow targeted exemptions for internal packages, registries without reliable timestamps, or recently adopted versions. The rollout began in familiar active repositories, supported by concise engineering communication, repository inventories created with Semgrep rules, and AI-assisted commands that updated manifests, lockfiles, CI configurations, Dockerfiles, and package-manager versions. As the effort expanded, Semgrep Agentic Workflows combined deterministic checks with Claude to automate repository discovery, exemption handling, lockfile validation, repairs, and pull-request creation, while Poetry support and registry-level exclusions broadened coverage. The experience also highlighted operational risks during package-manager migrations, the value of CI tests and fixing minor existing CI failures, and the benefit of archiving inactive repositories. Cooldowns are presented as a secure default rather than rigid governance, with exceptions available when developers need prompt access to security fixes, features, or internal releases.
Aug 31, 2026
2,481 words in the original blog post.
A comparative evaluation of Semgrep Multimodal, Claude Security with Mythos, and Codex Security on 275 manually reviewed IDOR labels found that Semgrep Multimodal achieved higher recall at 59.9% and F1 at 57.1%, while Mythos had higher precision at 80.1% but much lower recall at 13.9%; Codex Security recorded 11.3% recall and 17.7% F1. Across four repositories at identical revisions, Semgrep Multimodal reported 63 manually confirmed IDOR vulnerabilities not found by Mythos, with 40 recurring across three runs. The findings involved authenticated users accessing or modifying objects without authorization for the specific target object, often because checks covered a parent resource, object existence, workflow state, or authentication status rather than the child object or record used in the final operation. Semgrep Multimodal combines rule-based dataflow analysis, which traces caller-controlled identifiers through endpoints, service layers, and database operations, with AI reasoning intended to determine whether authorization checks apply to the same object as the sensitive action. The comparison emphasizes that IDOR detection requires assessing relationships among request parameters, authorization decisions, and final reads, writes, deletions, or workflow transitions, and that higher precision alone may leave many vulnerabilities undiscovered.
Aug 27, 2026
1,907 words in the original blog post.
A security benchmarking report evaluates Anthropic’s Mythos model in its Claude Security harness against open- and closed-weight models on 275 human-reviewed insecure direct object reference vulnerabilities across four codebases, using precision, recall, and F1 scores. Mythos achieved 80.0% precision but only 13.9% recall, identifying 20 of 144 confirmed vulnerabilities, placing fifth of 17 raw configurations for precision and fifteenth for recall; GLM 5.2 and GPT-5.6 Terra exceeded it on both metrics, while Claude Opus 5 achieved the highest raw recall at 38.2%. The comparison includes important methodological caveats, since Mythos was tested within a security-specific harness and evaluated with an internal LLM-based judge while other runs used deterministic matching, making model and harness effects difficult to isolate. Tests of GPT-5.6 Sol across three harnesses showed recall varying by 6.5 times, suggesting that prompting, context selection, and supporting tooling can influence results more than the underlying model. The report argues that security vendors should substantiate “Mythos-class” claims with reproducible recall measurements against labeled datasets, clarify evaluation methods and denominators, and continually rebenchmark because model performance and costs can change quickly.
Aug 25, 2026
1,612 words in the original blog post.
A Semgrep-led analysis of Hacker Summer Camp, spanning DEF CON, BSidesLV, and Black Hat, compared 2,295 talks from 2025 and 2026 using a consistent 27-category taxonomy, supplemented by transcripts from 556 recorded 2025 sessions containing 3.2 million words. It finds that the events serve different audiences and stages of security ideas: DEF CON emphasizes experimentation, BSidesLV translates emerging issues into practical professional threat models, and Black Hat focuses on difficult enterprise problems and commercial solutions. AI and machine learning discussions rose by roughly ten percentage points year over year and increasingly combined AI with established vulnerability and attack topics rather than replacing them, with AI-related offensive content growing more strongly than defensive content. Across talks, a prominent concern was improving the precision of AI-assisted security tools, especially reducing false positives through combinations of LLM reasoning, conventional scanners, rule-based analysis, and reliable validation methods. The analysis also notes that conference abstracts and titles may overstate certainty or hype compared with the more skeptical conclusions delivered in full presentations, while uneven recording practices mean in-person attendance remains important because much DEF CON content is never publicly released.
Aug 21, 2026
3,696 words in the original blog post.
On August 20, 2026, the widely used Rust crates arrayref and append-only-vec were compromised through malicious updates that added a dependency on proc-macro1, a typosquatted package impersonating the legitimate proc-macro2. Its build script executed automatically during Cargo builds, decoded obfuscated download URLs, retrieved platform-specific payloads for Linux, Windows, and macOS, and launched them as detached processes without requiring affected library functions to be called. A related package, proc-macro-en, appeared to impersonate arrayref maintainer droundy under the similar author name daveroundy, indicating broader campaign preparation. Organizations using proc-macro1 1.0.107, proc-macro-en 1.0.10, append-only-vec 0.1.9, or arrayref 0.3.10 were advised to rescan projects, review dependency and advisory records, and investigate identified command-and-control infrastructure and host artifacts.
Aug 20, 2026
260 words in the original blog post.
Semgrep evaluated Z.ai’s newly released GLM-5.3 and several Grok 4.6 variants on its IDOR vulnerability-detection benchmark, which tests models’ ability to identify missing authorization checks in real open-source code using precision, recall, F1 score, and cost per confirmed finding. Although Z.ai reports that GLM-5.3 has advanced cyber capabilities and strong CyberGym exploitation results, it achieved a 23.8% F1 score in this benchmark, roughly matching Claude Opus 4.8’s 23.6% while costing $0.15 rather than $1.04 per true positive, though it unexpectedly trailed its predecessor GLM-5.2 and will be retested for variance. Grok 4.6 Exacto scored 35.5% F1, approaching Kimi K3 and Claude Opus 4.7 performance at substantially lower cost, while Claude Opus 5 and GPT-5.6 Luna remained the leading models overall at 65.6% and 48.0% F1, respectively. Across most models, precision remained relatively high but recall was low, indicating that flagged vulnerabilities were often valid but that many real flaws went undetected; the results suggest that newer lower-cost models are increasingly competitive economically but have not yet reached current frontier performance in raw detection quality.
Aug 14, 2026
823 words in the original blog post.
Following the 2025 tj-actions/changed-files compromise, in which attackers redirected tagged releases to malicious commits, the author describes enabling GitHub’s full 40-character SHA pinning requirement across roughly 350 repositories to prevent workflows from executing mutable action tags or branches. The requirement applies not only to directly referenced actions but also to actions used transitively within composite actions and reusable workflows, making rollout more complex than simply converting tags such as `@v4` to commit SHAs. The effort began with small-scale testing that exposed issues involving internal actions referenced by branches and unpinned dependencies, then expanded through scripts and Semgrep Agentic Workflows to inventory repositories, generate remediation pull requests, and monitor failures. Pinact was used to convert tagged actions while Renovate was adopted to keep SHA-pinned internal and third-party actions updated with configured cooldowns and controlled major-version upgrades. The author also discusses GitHub limitations, including the lack of an evaluation mode, aggregate failure reporting, or exceptions under organization-wide enforcement, and notes complications involving reusable workflow behavior, CODEOWNERS, auto-merge, and abandoned third-party actions. The recommended rollout is to communicate the change, automatically protect new and inactive repositories, establish an update strategy, test and monitor enforcement repository by repository, resolve direct and transitive dependencies, and only then enable the setting organization-wide with monitoring and repository-level rollback options available.
Aug 12, 2026
3,727 words in the original blog post.
The "Worms are Back" ChainDrop campaign is an automated npm compromise observed on August 4, 2026, characterized by the republishing of legitimate packages under hijacked maintainer credentials. This approach led to 1,557 malicious versions appearing within about two hours. The compromised packages include obfuscated loader files integrated into the preinstall lifecycle hook, allowing code execution during dependency resolution. The second stage of this malware targets developer workstations and CI/CD runners to harvest credentials like npm authentication tokens, cloud provider credentials, SSH private keys, and CI secrets. The worm propagates by utilizing harvested npm tokens and employs an Ethereum dead-drop for command-and-control infrastructure, enabling operators to reassign infrastructure without hardcoding domains. Semgrep users are advised to scan projects for potential impacts, while indicators of compromise include specific file hashes, install hooks, and command-and-control resolution methods.
Aug 04, 2026
2,822 words in the original blog post.
The ongoing debate about measuring software development productivity, particularly using lines of code (LOC) as a metric, highlights its limitations and the evolving challenges faced with the rise of AI-assisted coding. Despite the historical reliance on LOC, experts like Frederick Brooks and Bill Gates have noted its inadequacies, especially when it ignores software requirements and penalizes efficient coding practices. The integration of AI in software development has dramatically increased coding output but also introduced security challenges, prompting a shift towards embedding security measures within AI tools rather than relying solely on traditional methods like continuous integration. As non-traditional developers increasingly generate code, ensuring secure practices becomes essential, with tools like Semgrep Guardian emerging to offer real-time security feedback. This approach reflects a broader trend of designing systems that mitigate risks associated with AI-generated code, emphasizing the need for integrated security solutions that consider both deterministic and probabilistic reasoning to address vulnerabilities effectively.
Aug 04, 2026
1,606 words in the original blog post.