JSFuck obfuscation tricked the Manus agent into running malicious code from a single email

Security researchers at Salt Labs managed to bypass the prompt-injection defenses of the AI agent Manus by hiding a malicious prompt in an email encoded with the JavaScript obfuscation technique JSFuck.

An opened envelope with tangled steel wire spilling out and a single red thread escaping — an illustration of hidden malicious code in an email.
Illustration
Gift article

JSFuck obfuscation tricked the Manus agent into running malicious code from a single email

Security researchers at Salt Labs managed to bypass the prompt-injection defenses of the AI agent Manus by hiding a malicious prompt in an email encoded with the JavaScript obfuscation technique JSFuck. The agent decoded and executed the code in its server-side environment — and the guardrail only raised the alarm after the code had already run. The researchers then established a reverse shell in the environment and found credentials and tokens for the user's connected third-party services. According to the researchers, the vulnerability has been closed, but the case shows why guardrails alone are not a defense for autonomous agents.

What happened

In early October 2026, several security outlets published coverage of Salt Labs' research: Salt Security's security team had found a way to bypass the guardrails of Manus, an agent platform that can connect to a user's email, cloud storage, and other services (Cyber Security Intelligence, TechRadar via MSN, SC Media).

According to Cyber Security Intelligence, which reproduces Salt Security's own account, the researchers conducted the study earlier this year. They first reported the issue to Manus, without receiving a response. They then submitted it through Meta's bug bounty program, where Meta, according to the reporting, confirmed and fixed the issue. As relayed in the TechRadar article, Salt Labs says the problem "has since been resolved and is no longer exploitable." It is worth noting that all of this rests on the researchers' own account — Manus itself has not confirmed the details in any of the covering sources.

The specific vulnerability has thus been patched, but the incident shows, as SC Media points out, an essential lesson for businesses deploying AI agents with broad access to third-party services.

The attack, step by step

First, what actually works in Manus' defense: According to TechRadar's description, the agent can detect when a prompt is hidden in data intended for analysis. But it does not ignore the prompts entirely — it alerts the owner when it finds a malicious one.

The researchers experienced this themselves. When they sent a test user an email with an explicit command, Manus' guardrail caught and flagged the obviously malicious instruction, Cyber Security Intelligence writes.

The next step was obfuscation. The researchers re-encoded the command with JSFuck, a JavaScript obfuscation technique which, according to the researchers as quoted via TechRadar/MSN, uses a limited character set and is therefore an "unusual JavaScript obfuscation technique" that is "rarely used in modern environments." The point of this kind of obfuscation is that the code appears as a meaningless string of characters rather than executable logic.

Now the agent no longer concluded that the content was a malicious prompt. It interpreted it, according to the researchers, as something to be decoded and rendered. But as the Salt Labs researchers said, in TechRadar's rendering: "The intent appeared to be content decoding and rendering. However, this effectively executed arbitrary JavaScript code in a server-side environment!"

That is the core flaw: the task "decode this" and the action "run this" merged. The obfuscation removed the signals the guardrail relied on to recognize malicious content, while the agent was still able to process — and in this case execute — the content. And the alert to the owner came only after the code had been run, according to the TechRadar article's account of the sequence of events.

From code execution to exposed tokens

Code execution in the agent's server-side environment is not just a technical breach of a security boundary — it provides a foothold. According to Cyber Security Intelligence, the researchers then established a reverse shell in the environment — a connection out of the environment that allows control from the outside — and identified credentials and tokens tied to the third-party services the user had connected to the agent.

Two caveats are important here. First, this was the researchers' demonstration, not a documented incident in which real user data was exfiltrated. Cyber Security Intelligence describes the string of exposed accounts — email, cloud storage, code repository, and similar — as a potential scenario, not a confirmed exfiltration. Second, the logical point is serious regardless: an agent that connects to third-party services on the user's behalf must store or be able to retrieve authorization for those services. An environment compromised through code execution can therefore give an attacker the keys to the services the agent can access — one email in, broad access out.

Why guardrails fail against obfuscated input

Manus' guardrail operates on the content as it appears: patterns of malicious instructions. This kind of defense resembles traditional detection of known attack signatures — effective against what it recognizes, blind to what is packaged in a form it does not recognize.

Prompt injection exploits precisely the fact that language models do not robustly distinguish between "instruction" and "data." A guardrail that inspects raw text can flag instructions in plain text. But obfuscated content shifts the problem: the agent's interpretation layer can decode the content along the way — because decoding is a legitimate task the agent is often asked to perform — and the model is thus confronted with the malicious instruction only after the input has already passed the outer defense layer.

The timing problem is at least as important as the recognition problem. The fact that Manus alerted the owner after execution means the guardrail functioned as an alarm, not as a lock. In an autonomous system where one action can trigger code execution, and the code execution can yield persistent access, an alert afterward is weak protection.

The researchers' warning: layer by layer, not a single barrier

Yaniv Balmas, head of research at Salt Security, summarizes the lesson as follows, as relayed by Cyber Security Intelligence: "Anyone designing an agentic system should build robust, layered defenses rather than relying on guardrails to provide all the protection, just as we learned to do with traditional services. As agent adoption grows, I have no doubt this will become one of the most common attack vectors we see."

The analogy to traditional security is real. Web and API security taught the industry that input validation alone is never enough: you add least privilege, isolation, authentication between services, monitoring, and limits on what a compromised environment can actually reach. Transposed to agentic systems, that means concrete measures such as running the agent's environment isolated with minimal privileges, keeping third-party tokens not readily accessible within the environment, requiring explicit approval for actions that enable code execution, and alerting before risky actions are performed — not after.

The tricky question of the responsibility line

One curiosity in this case is who actually owns the vulnerability. Salt Labs reported, according to Cyber Security Intelligence, first to Manus, without receiving a response, and then used Meta's bug bounty program — where Meta, according to the reporting, confirmed and fixed the issue. Two things are worth keeping apart: Meta's bug bounty program was used as the reporting channel, and Meta confirmed and addressed the issue there; the claim that Manus did not respond comes from Salt Labs itself, as relayed by Cyber Security Intelligence, and has not been confirmed by Manus.

For an industry where AI agents are built, operated, and integrated by different actors, this is not merely a procedural matter. Who users should report vulnerabilities to — and who actually patches them — becomes part of the security model of the entire agentic ecosystem.

Uncertainties and open questions

Everything known about this case comes from secondary sources reproducing Salt Security's research, and they likely rest on the same underlying report. Salt Labs' original technical report is not directly among the available sources, so the three covering outlets should not be read as independent confirmation of one another.

The timeline is also vague: Cyber Security Intelligence writes only that the research took place "earlier this year" and was reported to Manus without a response, without precise dates. The MSN/TechRadar article mentions only the disclosure through Meta's bug bounty program, while Cyber Security Intelligence adds the initial-contact-with-Manus step. The sequence cannot be fully verified.

And the biggest open question is the one SC Media points to: even though this specific technique is closed, there are likely others that are just as effective — an assessment Salt Labs itself makes in TechRadar's rendering. A guardrail that recognizes a pattern can be bypassed with a new wrapper around the same malicious content. That is why the point of this story is not that Manus was vulnerable, but that the defense design of agentic systems must assume any single layer can be broken — and ensure that the consequence of a breach is small.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.