GitHub's Taskflow Agent Found 24 Vulnerabilities in Android Apps, According to Its Own Case Study

GitHub Security Lab published a case study on 28 September 2026 in which Taskflow, an AI agent retooled for Android security auditing, found and reported 24 vulnerabilities in Android apps — among them silent location tracking in OsmAnd…

Illustration: a row of black miniature devices with one opened and exposed under an inspection lamp, symbolizing security auditing of Android apps.
Illustration
Gift article

GitHub's Taskflow Agent Found 24 Vulnerabilities in Android Apps, According to Its Own Case Study

GitHub Security Lab published a case study on 28 September 2026 in which Taskflow, an AI agent retooled for Android security auditing, found and reported 24 vulnerabilities in Android apps — among them silent location tracking in OsmAnd and account takeover in the Wikipedia app. But the figure and the details are GitHub's own claims as relayed through secondary sources, and the agent showed a clear bias: it was good at finding flaws, poor at judging how serious they were.

What happened

Kevin Stubbings, a researcher at GitHub Security Lab, built his own AI-driven audit flows — so-called "taskflows" — on top of the lab's open source Taskflow agent, and used them to find and report vulnerabilities in Android apps. According to Cybersecurity News, which relays GitHub Security Lab's own announcement, the total is 24 vulnerabilities; Help Net Security writes "more than 20." Both sources are secondary accounts of the case study, so the figure should be read as GitHub's own count — not as independently verified.

Among the findings, two cases stand out as illustrations of the kinds of vulnerabilities the agent can actually trace: a chain in the navigation app OsmAnd enabling hidden location tracking, and a two-step attack chain in the Wikipedia app for Android that can yield full account takeover.

How the agent works

Rather than scanning entire codebases with generic prompts, the Android adaptation consists of two phases. One taskflow identifies mobile entry points — exported activities, services, broadcast receivers and deeplinks. Another evaluates each entry point against Android-specific vulnerability classes: insecure intents, confused deputy problems, unsafe broadcasts, cross-app scripting and WebView risks.

The point of this split is that the agent does not merely summarize code, but follows concrete attack paths: first, "where can outside code get in?", then "what can an attacker get out of each entry point?" The results are collected in an SQLite table called audit_results.

Case 1: OsmAnd — route reconstruction without the user's knowledge

OsmAnd has more than 10 million downloads on the Play Store. In the app, an exported activity called MapActivity accepted intent extras that were supposed to be reserved for an internal channel. That means any app on the phone — with no permissions at all — could use these extras, which according to Cybersecurity News include silent_import, replace and export_type_list_key, to import malicious settings silently.

The attack chain looks like this: an unprivileged malicious app swaps OsmAnd's map tile source for an attacker-controlled server. The server then logs the coordinates of every map tile the victim loads. Since map tiles correspond to geographic areas, the attacker can reconstruct the victim's routes — all without the user noticing anything. Help Net Security describes it as location tracking without the user's knowledge, powered by the app's own map data.

Case 2: Wikipedia — account takeover after a single link tap

The Wikipedia app had two linked flaws. First, the hostname check in the deeplink handler used endsWith() instead of matching the full domain. This meant a wikipedia:// link could point to a lookalike domain such as evil-wikipedia.org and load it in the app's WebView — where the attacker's JavaScript runs in a context the app treats as trusted.

The second flaw sat in the cookie manager: the WebView could retrieve the victim's long-lived session cookies. Chained together, they give the attacker an authentication and session token that is valid across all Wikimedia projects — Wikipedia, Wikimedia Commons, Wikidata and Meta — after the victim taps a single malicious link.

Where the AI automation stops

That may be the most revealing finding in the case study: the model was better at finding flaws than at assessing how dangerous they were. According to Help Net Security, citing Stubbings, the model kept flagging low-risk findings even after being asked to stop, and it mishandled real impact in cases where a mitigating factor — such as internal storage that silently overrides attacker-controlled external storage — quietly neutralized what looked like a working exploit.

That is why GitHub warns against accepting AI findings without expert review: language models can recognize code patterns and relevant APIs, but they can also misjudge severity, overlook mitigating behavior or generate false positives. The researchers found that requiring the model to build a proof of concept significantly improves triage — but human testing remains necessary, and every finding needs a reviewer who knows mobile development.

Availability and cost

The Taskflow agent and the Android audit flows are open source and freely available. But running them is not free: a GitHub Copilot license is required, and an audit of a medium-sized codebase can take one to two hours and consume many premium model requests. No precise cost figures have been published.

Open questions

Part of the picture remains unclear. It is not confirmed whether all 24 vulnerabilities have been fixed, whether any of them have been assigned CVEs, or which apps beyond OsmAnd and Wikipedia are affected — none of the available sources list them. Moreover, the entire story rests on secondary accounts of GitHub Security Lab's own case study; none of the details have been independently verified.

It is nevertheless a concrete marker of a trend: AI-driven vulnerability research has moved from broad code review to tracing complete attack chains in large, widely used apps — with humans still at the end of the triage queue.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.