The safety your LLM vendor demoed is a peel, not a wall.

Unit 42 researchers just showed that refusal, the model’s canned “I can’t help with that,” lives in a thin neural band you do not operate. Their perturbation probing method jostles those activations and the no collapses. Harmful text comes out. If that refusal is the threat-protection in your cybersecurity design, you have a brochure where a control should be.

Overview graphic of perturbation probing used to test how fragile LLM safety refusal layers are
Unit 42 treats model refusal as a diagnostic target, not a guarantee you can put in a risk register.

Refusal Lives In A Thin Layer

You already knew jailbreaks exist. This work makes the geometry ugly. Safety fine-tuning parks the “no” in a small region of the network. A modest perturbation, not a theatrical roleplay prompt, knocks the model off that ridge and into helpful-and-harmful mode.

That is the opposite of a robust control.

Keep alignment work. Budget as if it will fail on contact. The Unit 42 write-up is explicit: because the internal safety layer is fragile, you need external, multi-layered security around the model. Defense in depth is the actual product. The weights are a component.

Your cyber security questionnaire still asks whether the model “refuses harmful requests.” Stop checking that box like it means something. Ask where policy is enforced, who can change tool allowlists, and what you log when the model tries anyway. If the answer is that the model knows better, you are buying a vibe.

You see the same move in products that are not models. WhatsApp shipped passkeys plus stronger two-factor options this week. Account takeover on a messenger your CFO already uses for “quick approvals” is an incident, not a lifestyle choice. The vendor finally bolted on real locks. Your program still has to assume that chat is an unmanaged channel until identity, device, and logging say otherwise.

WhatsApp logo after the app added passkey and two-factor account protections
Passkeys on a messenger help. They do not make executive chat a controlled system.

Cloudflare’s BotBase for operators offers bot owners a dashboard to declare how their agents use content, edit submissions, and track status. That helps polite crawlers. It is not threat detection. The operators who will fill out the form are not the ones you are hunting.

A firewall that trusts a model’s no is an open proxy with extra steps.

Brute-force against the login in front of your copilot portal is still cheaper than a research-grade jailbreak. If you hardened SSH and left the AI gateway on a shared password, you picked the wrong door to paint.

Cybersecurity Starts With Wrappers You Own

Stop treating the model as a policy engine. Treat it as an untrusted intern who types fast and agrees with whoever spoke last. Then put the real controls in software you run.

Do this now. Inventory every LLM, copilot, and agent that can call tools or see customer data. Name an owner. Kill anything that can reach mail, tickets, or file stores without an authenticating proxy. Move refusal out of the prompt. Allowlist tools, destinations, and data classes in the app layer so a jailbroken yes still gets a 403.

Turn on WhatsApp passkeys and two-step verification for executives, finance, and anyone on the incident response roster. Put it in writing. Personal phones used for work chats are in scope. If you cannot manage the app, you can still manage the people and the approval rules that make a hijacked thread expensive.

Keep going after the first week. Ship prompts, completions, and tool calls into the same threat detection pipeline you use for VPN and SaaS. Alert on jailbreak-shaped inputs, sudden tool-use spikes, and dumps that look like exfil. Give agents and chat backends a tight egress policy. No unrestricted browsing “so the model can look it up,” and no surprise outbound tunnels.

Red teams should steal the researchers’ idea. Do not only run public jailbreak lists. Perturb, rephrase, and inject retrieved documents that collide with policy. If your only test is a banned-word filter, you are testing the sticker, not the system.

Tabletop the boring failures. Your incident response plan needs “the model pasted a customer file into a ticket” and “an exec’s messenger got taken over,” not only ransomware. Security hardening of identity in front of every AI feature is non-negotiable: SSO, phishing-resistant MFA, short-lived tokens, lockouts. IP-level bans on those portals belong there for the same reason they belong on VPN. Attackers still try the front door.

Output-side filters still matter. Regex and classifiers are crude. They are also yours. A second scorer on completions is not magic, but it is a layer that does not live in the same thin band Unit 42 just mapped.

None of those jobs belong to a refusal vector a lab can poke over.

Patch the messenger. Constrain the agent. Log the tool calls this week.

Sources

Take Control of Your Server Security

Don't let brute-force attacks slow you down. Try IPBan Pro risk-free for 30 days.

Secure. Automated. Lightweight.

Stay up to date with the latest news, releases and more.

Take Control of Your Server Security

Don't let brute-force attacks slow you down. Try IPBan Pro risk-free for 30 days.

Secure. Automated. Lightweight.