Home / Companies / Galtea / Blog / October 2026

October 2026 Summaries

1 posts from Galtea

Filter
Month: Year:
Post Summaries Back to Blog
Galtea benchmarked TypeSafe’s English-trained Jev LLM judge on 1,049 human-labelled single-turn evaluation rows translated into Spanish and Catalan, comparing its results with a tuned GPT-5.2 judge across 10 metrics. Jev’s accuracy remained essentially unchanged for Spanish inputs and declined by only 1.5 points for Catalan when its rubric questions stayed in English, but translating the questions into Spanish or Catalan reduced average accuracy further and caused a 13-point drop in toxicity detection, partly because translated wording changed the interpretation of terms such as “dismissive.” Against GPT-5.2, Jev performed significantly better on factual accuracy and non-toxicity while producing statistically similar results on seven other metrics, although both judges struggled with jailbreak resilience. The study concludes that Jev can support Spanish and Catalan input evaluation effectively when prompts remain in English, while noting that its suitability for multi-turn conversations, agent traces, and complex evaluation templates requires further research.
Oct 03, 2026 2,018 words in the original blog post.