GitHub Turned Its Audit Prompts Into an Open-Source Agent — It Found 24 Android Vulnerabilities
On 28 September 2026, GitHub Security Lab researcher Kevin Stubbings published the rare defensive-AI story with a body count: custom audit workflows ("taskflows") built on the lab's open-source Taskflow Agent have found and reported 24 vulnerabilities in Android applications — and you can run the same harness against your own repos. Two disclosed examples show the range: covert location tracking in OsmAnd (10M+ Play Store installs) via an exported activity, and account takeover in the Wikipedia app via a deeplink parsing bug, both discovered by prompting strategies tuned for mobile attack surface rather than by a bigger model.
The method: split the attack surface before you prompt
Stubbings' key move was not a new model but task decomposition aimed at Android's vulnerability classes. A new gather_mobile_entry_point_info.yaml taskflow separates mobile entry points from web/desktop ones, so the agent reasons about the right attack surface even in mixed repos. An edited classify_application_local.yaml then pins a checklist of mobile-specific bug classes — confused deputy, insecure broadcasts, intent-based flaws — to each entry point, compensating for the fact that mobile vuln classes are underrepresented in training data and LLMs are non-deterministic. Strict prompts plus repeated runs catch the obvious bugs; a broad prompt layered over them lets the model be creative. It is the same lesson as the authorized-offensive-agent work we covered in Wiz's Scan-for-Good: the scaffolding around the model is the product.
OsmAnd: any app, no permissions, your location history
The most instructive of three OsmAnd findings starts with an exported MapActivity that handles settings files and deeplinks. When opening settings it honors intent extras — settings_version, silent_import, replace, export_type_list_key — that were only ever meant to arrive from an internal AIDL service. But Android provides no mechanism to restrict which extras an external caller attaches, so any app on the phone can fire an intent at the exported activity carrying silentImport=true (no notification, no user confirmation) and replace=true (overwrite, don't merge). handleOsmAndSettingsImport complies.
From there the attacker rewrites the map tile URL template to point at their own server. Every tile the victim loads leaks its exact x/y/zoom coordinates; the attacker's backend serves the genuine OpenStreetMap tiles back, so nothing visibly changes while routes are reconstructed tile by tile — the write-up demonstrates per-tile coordinates resolving to street-level locations and full route origin/destination pairs exfiltrated the same way. No permissions on the malicious app. No user-visible symptom. Ten million installs of implicit trust in an exported component.
Wikipedia: one tapped link, every Wikimedia session
The Wikipedia app registers the wikipedia:// deeplink scheme — and its hostname check used endsWith() instead of matching the full domain, so a link pointing at a lookalike like evil-wikipedia.org loads inside the app's WebView, running attacker JavaScript in a context the app treats as trusted. A second flawed check in the app's cookie manager then let that WebView pull the victim's long-lived session cookies — tokens valid across every Wikimedia project. One tapped link, total cross-project session theft. It is a textbook reminder that mobile deeplinks are unauthenticated entry points wearing a trusted-UI costume.
Honest limits: better hunter than judge
To Stubbings' credit, the write-up reports the failure modes, not just the trophies. The model kept flagging low-severity issues after being told not to, and misjudged real-world impact where a mitigating factor — e.g., internal storage silently overriding attacker-controlled external storage — defused an apparent exploit. Every finding still requires a human reviewer who understands mobile apps before it ships anywhere. That mirrors the broader pattern in agentic security tooling: recall is cheap, precision and impact-judgment are the scarce resources, and the human is the precision layer.
What to do today
- Run the harness against your own Android code. The taskflows are open source in the
seclab-taskflowsrepo: open a codespace, run./scripts/audit/run_mobile.sh myorg/myrepo, wait an hour or two for a medium repo, and triage rows flagged in theaudit_resultsSQLite table. Budget for a GitHub Copilot license and heavy premium-request consumption. - Audit every exported component for assumed-private extras. The OsmAnd pattern — extras that are safe only because "only our code sends them" — is endemic. If an activity, service, or receiver is exported, every extra it reads is attacker-controlled by definition.
- Grep deeplink handlers for suffix matching. Any
endsWith/contains-style host validation on custom schemes is a lookalike-domain bug waiting for a tap. Match full domains, and scope WebView cookie access to authenticated origins. - Keep the human in the triage loop by design. Deploy agent-found findings as leads with exploitability notes, not as auto-filed CVEs — and staff the reviewer role with someone who knows the platform's silent mitigations.
- Watch the advisories page, not just the blog. Stubbings notes further disclosures will land on the lab's advisories page as they publish; the 24 count is a floor, not a ceiling.
Sources:
- GitHub Blog, Kevin Stubbings — "How we found 24 Android vulnerabilities using our open source AI security agent" (28 September 2026: Taskflow Agent, gather_mobile_entry_point_info.yaml and classify_application_local.yaml taskflows, 24 vulnerabilities found and reported, OsmAnd exported MapActivity intent-extras location tracking, Wikipedia deeplink account takeover, run_mobile.sh reproduction, Copilot license and premium-request costs)
- Help Net Security, Anamarija Pogorelec — "GitHub's AI agent found 24 Android app vulnerabilities" (29 September 2026: OsmAnd tile-source swap and route reconstruction detail, Wikipedia endsWith lookalike-domain plus cookie-manager session theft across Wikimedia projects, severity-judgment limits, human-reviewer requirement)
- CybersecurityNews — "GitHub AI Security Agent Finds 24 Android Vulnerabilities Including Account Takeover Flaws" (29 September 2026: disclosure scope confirmation)