ai-security-research-radar

mcp
Security Audit
Warn
Health Warn
  • No license — Repository has no license file
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 5 GitHub stars
Code Pass
  • Code scan — Scanned 1 files during light audit, no dangerous patterns found
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

AI Security Research Radar — AI for security and security for AI

README.md

AI Radar

trends accelerating watchlist updated

Autonomous tracker of the offensive AI-security frontier — AI for offense and attacks against AI — for a security researcher; generated from TRENDS.md.

Since last scan (2026-10-02) — weekly W40 recalibration:

  • 🚀 Agent authorization-integrity goes accelerating: agent authorization & identity integrity promoted emerging→accelerating — 9 pairwise-independent groups across 7 facets (memory-laundering, self-issued-auth, revocation-enforcement, effect-closure, HITL-approval hijack, orchestration-identity, coding-agent-harness approval laundering) on a sustained 4-week cadence; confidence held medium (still all-academic, no in-the-wild case yet).
  • 🖼️ Provenance-attack trend broadened to image / audio / text: provenance & watermark attacks title broadened across modalities (W39 proposal applied; the persisting signal was a 2nd audio-removal group, DeMark).
  • 🛰️ New primary feed promoted: GitHub Security Lab → swept every run — its open-source Taskflow Agent AI-vuln-discovery series (24 Android bugs + scan-framework + AI-triage) cleared the ≥2-on-axis bar.
  • 🧹 Housekeeping: 6 stale run-log husks (09-14→09-19) burned down, 7 study-shelf picks pruned (>30d); slopsquatting held dormant (archive ~10-08 if silent — "in-the-wild" slopsquat cases found this week are old Feb-2026 recycling); in-the-wild AI-offense held accelerating, next dormancy re-check ~10-10.

Trends

🌱 0 · 📈 7 · 🚀 8 · 🌊 0 · 🏔 0 · 📉 0 · 💤 1

trend stage latest signal
Attacks on LLM-agent stack: MCP, skills, supply chain 🚀 accelerating 2026-09-30
AI-security tooling unreliable: scanners, guards, judges 🚀 accelerating 2026-09-30
Agent authorization & identity integrity 🚀 accelerating 2026-09-30
LLM/agentic vuln discovery, repair & AI-written code 🚀 accelerating 2026-09-29
Mechanistic basis of jailbreaks: refusal & harmfulness directions 🚀 accelerating 2026-09-29
RAG knowledge/document poisoning 🚀 accelerating 2026-09-28
Adversarial trigger implantation & backdoor attacks 🚀 accelerating 2026-09-21
In-the-wild AI-for-offense: LLM malware dev & C2 🚀 accelerating 2026-09-09
Undetectable covert channels & agent collusion 📈 emerging 2026-10-01
Provenance & watermark attacks (image/audio/text) 📈 emerging 2026-09-27
Physical-channel PI on embodied & wearable AI 📈 emerging 2026-09-25
Automated red-teaming of AI agents 📈 emerging 2026-09-25
Economic/availability DoS on LLM systems 📈 emerging 2026-09-25
Model extraction, distillation & fingerprinting 📈 emerging 2026-09-18
Self-evolving-agent skill poisoning 📈 emerging 2026-09-15
Weaponized LLM hallucination (slopsquatting supply chain) 💤 dormant 2026-07-14

🛠️ Tools & releases

PyPI registry-SEARCH tool candidates verified this weekly: Basilisk DROPPED (the "AI red-team framework" tvly-snippet collides with an unrelated 2015 basilisk NoSQL mapper on PyPI — no real on-axis package under that name); ptai confirmed real/recent (v1.4.1, 2026-09-12, "AI pentesting that proves its findings") but single-author with no adoption signal → carried below the notability bar; akio/pentest-ai stale/early → low-priority carry. Watched-repo releases this scan: all unchanged except giskard 3.0.0 → 3.0.1 (patch). Black Hat Arsenal / DEF CON Demo Labs off-season (next: Aug 2027). The current verified on-axis tool set:

  • NVIDIA/garak — the LLM vulnerability scanner; v0.17.0 (2026-09-09).
  • promptfoo/promptfoo — prompt/agent/RAG red-teaming & pentesting; v0.123.1 (2026-09-18).
  • microsoft/PyRIT — Python Risk Identification Tool for generative AI; v1.1.0 (2026-09-04).
  • Tencent/AI-Infra-Guard — full-stack AI red-team platform: Agent-Scan, MCP-Scan, Skill-Scan (SARIF 2.1.0), jailbreak eval (26+ methods); v4.6.0 (2026-08-26).
  • Giskard-AI/giskard — evals, red-teaming & test generation for LLM/agentic systems; v3.0.1 (2026-10).
  • confident-ai/deepteam — framework to red-team LLMs and AI agents; v1.0.9 (latest on PyPI).
  • aliasrobotics/cai — Cybersecurity AI (CAI): an offensive-AI-sec agent framework (surfaced via tool-discovery, established repo).
  • FuzzingLabs/mcp-security-hub — a Dockerized collection of 38 offensive-security MCP servers / 300+ tools (Nmap, Ghidra, Nuclei, SQLMap, Hashcat, …).

Worth studying


Community pulse

Unverified sentiment (Phase-3 intake, link-only) — never trend evidence.

  • Prompt-injection firewalls for agent tool calls surfaced again on Hacker News (an open-source PI-firewall for AI APIs / tool calls) — defense-tooling discourse adjacent to AI-security tooling unreliable; intake-only.
  • An intentionally-vulnerable web app for security training ("WattzGOAT") posted on HN — training-target discourse, adjacent to the vulnerable-by-design AI-app targets tracked in tool-discovery; intake-only.
  • Jailbreak-corpora stream (automated-redteaming eval teaser, Indic-Jailbreak-Bench, JailbreakDB) surfaced on the Hugging Face hub — intake artifacts on the jailbreak axis, no promotion.
  • Through the week: general PI / MCP / jailbreak recirculation only — no offensive-AI earthquake, no new untracked-topic vocabulary.

📄 TRENDS.md · 👁 watchlist (~13) · 🗂 reports/ → 2026-10-02 · 📅 weekly: 2026-W40 · 📘 AGENTS.md · 🌐 SOURCES.md

Reviews (0)

No results found