tech 🔬 Grok 4.6 shows strong biosecurity skills
LatchBio analyzed Grok 4.6's biological capabilities and biosecurity performance using specialized benchmarks. The model proved the strongest in refusing disguised and hazardous tasks while still handling routine biological work. Specifically, Grok 4.6 scored above 50% on both the BioSecBench-Refusal and routine compliance measures. On the BioSecBench-Surveillance test, it achieved a 53.5% success rate, placing it between Opus 5 and GPT-5.6 Sol. These results suggest Grok 4.6 is well-calibrated for both helpful science and critical safety measures.