Home / Companies / Galtea / Blog / Post Details
Content Deep Dive

Does an AI judge need to speak your customer's language? | Galtea Blog

Blog post from Galtea

Post Details
Company
Date Published
Author
-
Word Count
2,018
Company Posts That Month
1
Language
English
Hacker News Points
-
Post removed?
No
Summary

Galtea benchmarked TypeSafe’s English-trained Jev LLM judge on 1,049 human-labelled single-turn evaluation rows translated into Spanish and Catalan, comparing its results with a tuned GPT-5.2 judge across 10 metrics. Jev’s accuracy remained essentially unchanged for Spanish inputs and declined by only 1.5 points for Catalan when its rubric questions stayed in English, but translating the questions into Spanish or Catalan reduced average accuracy further and caused a 13-point drop in toxicity detection, partly because translated wording changed the interpretation of terms such as “dismissive.” Against GPT-5.2, Jev performed significantly better on factual accuracy and non-toxicity while producing statistically similar results on seven other metrics, although both judges struggled with jailbreak resilience. The study concludes that Jev can support Spanish and Catalan input evaluation effectively when prompts remain in English, while noting that its suitability for multi-turn conversations, agent traces, and complex evaluation templates requires further research.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Jev 50 No monthly metrics for this publish month.
LLM 10 No monthly metrics for this publish month.
AI Guardrails 3 No monthly metrics for this publish month.
AI Agents 2 No monthly metrics for this publish month.
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.