GitHub's security agent found roughly 24 vulnerabilities in Android apps
GitHub Security Lab disclosed on September 28, 2026 that the company's open-source security agent, the Taskflow Agent, helped identify roughly 24 vulnerabilities in Android apps — among them a flaw in the OsmAnd map app that could record the victim's movements via map tiles, and a vulnerability chain in Wikipedia's Android app that could enable account takeover. The count varies between sources: while Cyberpress, iSec News and TechTimes give 24, Help Net Security and Let's Data Science write "more than 20." Common to the reports is that the maintainers concerned received the notifications, and that the report did not identify exploitation in real attacks, according to iSec News.
The story is as much about the limitations as about the findings. Both GitHub and the researcher behind the project, Kevin Stubbings, are open about the model repeatedly misjudging severity and overlooking mitigating controls. At the same time, the workflow has now been published as open source and can be adopted by anyone with a GitHub Copilot license.
The two main findings
The most striking finding lies in OsmAnd, an open-source navigation app. The app's exported MapActivity accepted attacker-controlled intent extras, which gave an app with no permissions the ability to change the URL from which the map tiles are fetched. By redirecting the map tile requests to an attacker-controlled server while the app displays legitimate map imagery, an attacker could record the tile coordinates the victim's device requests, according to the GitHub finding as reported by Cyberpress. Since map tiles correspond to geographic areas, this could — as an extension of the analysis — make it possible to infer the victim's position and routes without any location permission having been granted. Cyberpress characterizes the finding as hidden location tracking.
The second finding concerns Wikipedia's official Android app. The deep-link handler used an endsWith() check for hostnames, a classic case of weak validation that allows domains that merely "end with" the correct name to slip through — for example, lookalike domains controlled by an attacker. These could be loaded in the app's WebView. Chained with a weak check in the cookie handling, the exposure could, according to the GitHub finding as reported by Cyberpress, enable account takeover across Wikimedia services, including Wikipedia, Commons, Wikidata and Meta. TechTimes describes the chain in stronger terms, with session tokens valid across all Wikimedia projects, but that characterization is a single outlet's retelling and should be read accordingly.
How the agent works
Taskflow Agent is built on the OpenAI Agents SDK and supports three different AI backends: OpenAI Agents (the default), GitHub Copilot SDK and Anthropic's SDK, according to TechTimes. A recurring feature of the architecture is the division between deterministic and reasoning tasks: Model Context Protocol (MCP) handles the predictable tool operations — file retrieval, API calls and YAML parsing — while the language model's reasoning is reserved for code understanding and vulnerability reasoning. The results are stored in SQLite.
What was decisive for the Android findings, however, was not the framework but the taskflows Stubbings built for mobile. One of them, gather_mobile_entry_point_info, maps, according to TechTimes, attack surfaces that are unique to Android — entry points that differ from their web and desktop counterparts. The taskflows explicitly instruct the model to look for intent-based bugs, such as "confused deputy" problems and insecure broadcasts — categories that generic security prompts typically overlook, according to Help Net Security. It is this mobile specialization that explains why the agent found precisely intent-dependent bugs like the OsmAnd tile finding and the Wikipedia deep link.
Where the judgment falls short
Both GitHub and Stubbings are clear that the agent's results cannot be taken as proof. GitHub warned that models can find low-impact issues, overestimate severity, and produce false positives when they overlook mitigating controls or fail to understand runtime behavior, according to Cyberpress.
Stubbings described two recurring failure modes, as reported by Help Net Security: the model kept flagging low-severity issues even after being told not to, and it misjudged the real impact in cases where a mitigating factor silently neutralized what looked like a working exploit. As a concrete example, internal storage is mentioned as silently overriding attacker-controlled external storage — which makes an apparently usable attack infeasible, but which the model failed to catch.
For security teams, this means the agent works best as a broadly deployed detector that proposes candidates, not as a judge of severity. Every finding must be verified manually before it is reported to maintainers.
What it costs
The entry ticket is a GitHub Copilot license, and running it consumes requests against premium models. The time required is given as one to two hours on a medium-sized codebase, according to GitHub as reported by Cyberpress. That places the workflow in a middle zone: too heavy for continuous runs in CI across all repositories, but affordable enough for periodic security reviews of apps with real users.
Open questions
Several matters remain unresolved in the available reporting. It is not confirmed whether the reported bugs have been patched, or whether any of them have received CVE identifiers — TechTimes' headline refers to "24 CVEs," but no CVE IDs appear in the source material, and that framing should be treated with caution. The vulnerability count varies, as noted, between "24" and "more than 20" across the sources. The reporting is based on secondary coverage of GitHub Security Lab's disclosure on September 28, 2026, which Let's Data Science cites github.blog as the source for; the details stand or fall with how faithfully the various outlets relay that original disclosure.
The biggest unanswered question, however, is generalizability: roughly 24 findings across Android apps shows that the method works for intent-driven bugs, and the OsmAnd and Wikipedia findings are the two most detailed examples described. But it remains to be seen whether other teams achieve a similar hit rate — and whether the benefit holds up once the manual verification of all false positives is factored in.
Sources: Cyberpress (de6cd65d), Help Net Security (4f54f795), Let's Data Science / iSec News (e58a4f8e), TechTimes (1187a4fa).

