Reviewing Malware with LLMs: OpenAI vs. Vertex AI
Blog post from Endor Labs
Large Language Models (LLMs) are evaluated for their ability to assess the risk of potentially malicious code snippets, with a focus on the comparison between OpenAI's models and Google's Vertex AI. The assessment process involves using LLMs to assign a risk score on a scale of 0 to 9, rather than a binary classification, which helps in more nuanced evaluations. Improvements in the LLM-assisted review process include removing comments to reduce the risk of prompt injection and expanding context size. In the comparison, OpenAI's gpt-3.5-turbo and Google's text-bison often agree on risk scores, but diverge in cases where code is obfuscated. OpenAI's gpt-4 model is noted to provide superior explanations and risk evaluations for non-obfuscated code. The blog also discusses the manageable risk of prompt injection in this context, due to the requirement for code to adhere to syntactic rules, which allows for pre-processing steps like removing comments to mitigate risks.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 11 | 1,948 | 218 | 98 | +23% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.