qualiow-exploratory-testing-skills
Health Warn
- License — License: MIT
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 7 GitHub stars
Code Fail
- execSync — Synchronous shell command execution in bin/mobile-cli.mjs
- exec() — Shell command execution in bin/mobile-cli.mjs
- process.env — Environment variable access in bin/mobile-cli.mjs
- process.env — Environment variable access in bin/wkeval.mjs
Permissions Pass
- Permissions — No dangerous permissions requested
No AI report is available for this listing yet.
wico - AI-powered exploratory testing skills for Claude Code using Playwright CLI
@qualiow/exploratory-testing
qualiow -- AI-powered exploratory testing skills using Claude Code + Playwright CLI.
Claude acts as a Principal QA Engineer: it explores web applications systematically, finds bugs, and generates structured reports. No test scripts, no framework -- just AI reasoning, browser commands, and QA heuristics.
How to Install This Project
Prerequisites
- Node.js >= 18
- Playwright CLI (
npm install @playwright/cli) - Claude Code (for running exploratory sessions)
Quick Start
# install globally
npm install -g @qualiow/exploratory-testing
# initialize in your project
qualiow init
# run your first session
/qa-explore https://your-app.com
Local Development
# clone the repository
git clone https://github.com/willcoliveira/qualiow-exploratory-testing-skills.git
cd qualiow-exploratory-testing-skills
# install dependencies
npm install
# build
npm run build
# run tests
npm test
How to Run the CLI
| Command | Description |
|---|---|
qualiow init |
Initialize qualiow in the current project |
qualiow explore [url] |
Run an exploratory testing session |
qualiow validate |
Validate all configuration files (targets, knowledge, domains) |
qualiow report |
Generate a report from an existing session |
qualiow list <type> |
List sessions, knowledge, targets, or domains |
qualiow gather |
Gather and analyze requirements for testing context |
How to Run Exploratory Sessions
Full session (45 minutes)
# explore a public site by URL
/qa-explore https://testers.ai/testing/
# explore using a saved target configuration
/qa-explore --target company-staging
Quick focused session (15 minutes)
# check a single page or feature
/qa-explore-quick https://app.example.com/checkout
Mobile session (simulator/emulator)
Runs a full exploratory session on an iOS Simulator or Android Emulator — either a
native app (the installed app is the system under test) or a mobile web app in the
real device browser (iOS Simulator Safari / Android Emulator Chrome — the genuine
engines, which Playwright cannot drive). The mode comes from the target config.
1. One-time machine setup
# installs everything scriptable: Maestro, JDK 17, Android SDK + Google-Play AVD
# (qa_pixel_api35), ios-webkit-debug-proxy, and the qa-iphone simulator device
scripts/setup-mobile.sh # add --android or --ios to scope; --yes for non-interactive
Two iOS steps are genuinely manual (the script prints them if needed):
# a) install Xcode from the App Store, then:
sudo xcode-select -s /Applications/Xcode.app/Contents/Developer
sudo xcodebuild -license accept
# b) if no iOS runtime is installed:
xcodebuild -downloadPlatform iOS
# then re-run scripts/setup-mobile.sh --ios (creates the qa-iphone simulator)
2. Verify readiness (read-only, run before every first session on a machine)
scripts/doctor-mobile.sh # or --ios / --android; fix each ✗ until READY
3. Create a target and run
Copy a template from data/targets/, rename it, and point it at your app:
| Template | What it tests | Driver |
|---|---|---|
_example-native-mobile.yml |
An installed iOS/Android app (native mode) | mobile-cli (Maestro) |
_example-sim-ios-safari.yml |
A mobile web app in REAL iOS Simulator Safari | mobile-cli (Maestro) |
_example-sim-android-chrome.yml |
A mobile web app in REAL Android Emulator Chrome | mobile-cli (Maestro) |
_example-mobile-emulation.yml |
A mobile web app via Playwright device emulation (no simulator) | /qa-explore (playwright-cli) |
/qa-explore-mobile --target my-mobile-target # real sim/emulator (native or web)
/qa-explore --target my-emulation-target # Playwright mobile emulation
4. Driving the device manually (what the skill does under the hood)
# Android: boot the standard AVD # iOS: boot the standard simulator
~/Library/Android/sdk/emulator/emulator \
-avd qa_pixel_api35 -no-snapshot & xcrun simctl boot qa-iphone && open -a Simulator
export MOBILE_CLI_STATE=/tmp/my-target-state.json # isolate this session's device/app state
bin/mcli set-device emulator-5554 # or the iOS sim UDID (xcrun simctl list devices)
bin/mcli set-app com.android.chrome # or com.apple.mobilesafari, or your app's id
bin/mcli launch
bin/mcli open-url https://staging.m.example.com # web mode only
bin/mcli snapshot # page/app as tappable refs e1, e2, ...
bin/mcli fill e4 my-username
bin/mcli snapshot # ALWAYS re-snapshot after the keyboard appears
bin/mcli click e7
bin/mcli screenshot /tmp/evidence.png
bin/mcli logs --errors --since 60
bin/mcli --help # full command surface
iOS Safari web targets additionally get real JS/DOM access via the WebKit bridge:
bin/wk-ios 'document.title' # eval JS in the sim's Safari
bin/wk-ios 'document.querySelectorAll(".cart_item").length' # exact-DOM assertions
bin/wk-ios --stop # stop the proxy after the session
Notes that save time:
- Never
relaunch-cleana web target — it wipes the logged-in browser profile. - SSO/MFA logins are done once, by hand, on the device; the browser profile persists
the session. There is no transferablestorage_stateon a real device browser. - Use
bin/wadbfor raw adb commands (logcat,pm,am) — it setsANDROID_HOMEfor you. - Full setup guide + troubleshooting:
docs/MOBILE-SETUP.md.
Generate or regenerate a report
# regenerate report for an existing session
/qa-explore-report output/sessions/2026-03-28-parabank
Add knowledge to the base
# add new heuristics, techniques, or documentation
/qa-knowledge-add
Set up a new target
# configure authentication, scope, and domain for a target app
/qa-target-setup
How to Verify Backend Acceptance Criteria
Not every acceptance criterion is visible in a browser. /qa-verify-backend covers tickets
whose ACs live below the UI — tables and streams, queue consumers, Lambda triggers, IAM
policies, webhooks, IaC, and the service's own HTTP endpoints — and produces an AC
traceability matrix instead of a pass/fail claim.
# Copy the template, point it at your service, then verify a ticket branch
cp data/targets/_example-backend.yml data/targets/local-my-service.yml
/qa-verify-backend --target local-my-service --context output/context/TICKET-123-context.md
/qa-verify-backend --target local-my-service --static-only # no cloud credentials
/qa-verify-backend --target local-my-service --api-only --parity local-my-service-staging
It runs four lanes:
| Lane | What it does |
|---|---|
| Static | Reads the implementation branch against each AC with git show (never checks out) — spec drift, scope creep, failure paths, identity propagation, producer/consumer contract breaks |
| Live | Read-only aws-cli probes, one per AC, raw output saved as evidence. Never mutates; hard stop on production |
| API | Calls the service's own endpoints from the already-authenticated page, with a written case matrix covering everything the browser's guards prevent. Read-only, limited to the target's api.probe_allowlist |
| End-to-end | Drives the real write path in a non-production environment, then re-probes the data layer — including same-tick writes, deletes, bulk saves, a second identity, and DLQ depth |
Each AC gets a verdict with cited evidence — PASS, PARTIAL, FAIL, BLOCKED,NOT-REACHABLE or UNVERIFIABLE — and the verdict names the mode it was reached by and the
environment it holds in. PASS (direct request, raw response attached) and PASS (read the diff) are different claims, and a verdict backed only by a code reading is UNVERIFIABLE,
never PASS. BLOCKED is a first-class outcome: when credentials are missing, the probe
commands are written out ready to run rather than the verdict being inferred from source.
Two things it will not let you get away with
Believing a green result from an environment that does not run your code. Before the first
probe, the skill fingerprints every component of the request path: which build is deployed,
whether the changed path is selected here, and whether the commit under test is genuinely an
ancestor of what is running. One identical build can hold two implementations of the same
feature with a flag choosing between them — same build, different behaviour ⇒ configuration,
not deploy lag. An environment that does not run the change gets NOT-REACHABLE.
Recording a client-side guard as the endpoint's behaviour. "The search box refuses to fire
when empty" describes the browser. Every other client — mobile, partner integration, script —
sends that request, and it is frequently the one that fails. When a ticket has both a screen
and an endpoint, the same cases run at both surfaces and each finding is sorted into both,
API only (a real defect the guard is hiding), UI only (the client invents or masks
behaviour the service does not have) or neither.
The API lane needs no token handling: the request runs inside the page that is already logged
in, so it carries the same session cookie, CSRF token and interceptors as the UI — which also
means it works with SSO and MFA that no scripted login can pass. Credentials stay in the
browser profile under .auth/ and never enter a script, a report or the repository.
Correct shape is not a correct answer
Every derived value — a percentage, total, ratio, delta or aggregate — is recomputed from the
raw figures in the same response, with the formula taken from the specification rather than from
the code under test, and with cases chosen to stress sign, zero, scale and cardinality. Then the
report states the limitation plainly: when both sides of the check come from one payload, the
derivation is verified and the inputs are not, so the independent oracle that would close
the gap is named along with whether it was run.
The payload is also held next to the screen, because a correct response can still reach the user
as a wrong number — a formatter that guesses what a value is, a unit applied twice, a truncated
figure presented as a total. That defect is invisible from either surface alone, and it usually
belongs to a different change than the one under test.
Findings with no acceptance criterion become a spec
Most of what an API probe turns up has nothing to be filed against, so it turns into an argument
rather than a fix. data/templates/expected-behaviour.md converts a pile of observations into
one reviewable decision: observed against expected, grouped by cause, with the decisions the fix
forces made explicit — reject rather than clamp, an error rather than a silent zero, validation
at the layer that covers every implementation — ranked by what real users can reach today, and
closed with a plain-English reply for whoever decides to fund the work.
Verifying a release rather than a ticket
.claude/skills/qa-verify-backend/references/release-readiness.md changes the shape of the
session: the deployment table first
for every ticket, so a ticket whose backend is not in the build is marked not testable here
rather than tested against a UI that will render convincing nonsense; a result vocabulary that
keeps not tested visible; a coverage map where 🔍 code-verified only is marked asUNVERIFIABLE rather than green, with one sentence naming the untested item that carries the
most risk; carry-over defects in their own section; and a disposition with the condition that
would reverse it.
Full guide: docs/BACKEND-VERIFICATION.md. Safety rules (read-only discipline, the
production hard stop, redaction): .claude/skills/qa-verify-backend/references/safety-rules.md.
How to Generate Reports
qualiow produces three report formats from session data.
HTML report
Self-contained dark-mode HTML with severity cards, bug details, coverage map, and recommendations. Print-friendly.
# from TypeScript
import { generateHtmlReport } from '@qualiow/exploratory-testing';
const html = await generateHtmlReport('output/sessions/2026-03-28-parabank');
JSON report
Structured JSON for programmatic consumption -- dashboards, CI integrations, metrics tracking.
# from TypeScript
import { generateJsonReport } from '@qualiow/exploratory-testing';
const data = await generateJsonReport('output/sessions/2026-03-28-parabank');
Jira CSV export
CSV importable via Jira bulk import. Maps severity to Jira priority (Critical -> Blocker, High -> Critical, Medium -> Major, Low -> Minor).
# from TypeScript
import { generateJiraExport } from '@qualiow/exploratory-testing';
const csv = await generateJiraExport('output/sessions/2026-03-28-parabank');
How to Use Skills in Claude Code
| Skill | Purpose |
|---|---|
/qa-explore |
Full exploratory testing session (45 min) |
/qa-explore-mobile |
Exploratory session on a simulator/emulator — native apps or mobile web in the real device browser (iOS Safari / Android Chrome) |
/qa-explore-quick |
Quick focused session on a single page or feature (15 min) |
/qa-verify-backend |
Verify backend/API/infra acceptance criteria with no UI surface — branch review, read-only cloud probes, direct API probes, AC traceability matrix |
/qa-explore-report |
Generate or regenerate report from existing session |
/qa-explore-feedback |
Post-session feedback capture (false positives, missed bugs) |
/qa-explore-cleanup |
Session cleanup and archival |
/qa-knowledge-add |
Add new heuristics, techniques, docs to knowledge base |
/qa-knowledge-list |
Browse the knowledge base |
/qa-target-setup |
Configure a new target application (auth, scope, domain) |
/qa-gather |
Gather and analyze requirements from files, URLs, or text |
Project Structure
qualiow-exploratory-testing-skills/
src/
cli/
index.ts # CLI entry point (commander)
commands/ # CLI command handlers
formatters/
html-report.ts # standalone dark-mode HTML report
json-report.ts # structured JSON report
jira-export.ts # Jira-importable CSV
index.ts # re-exports
schemas/ # Zod validation schemas
types/ # TypeScript type definitions
utils/
metrics.ts # session metrics (JSONL append/read)
parse-session.ts # parse session dirs into structured data
redact.ts # credential and secret redaction
validate.ts # YAML config validation
index.ts # public API re-exports
data/
knowledge/ # YAML knowledge base (heuristics, techniques)
templates/ # bug report, session report, charter templates
targets/ # target application configs
domains/ # domain-specific testing focus
security/ # security policy
output/
sessions/ # session outputs (reports, bugs, screenshots)
tests/
unit/ # unit tests (vitest)
fixtures/ # test fixture data
benchmarks/ # performance benchmarks
docs/
GETTING-STARTED.md # detailed setup and usage guide
BACKEND-VERIFICATION.md # verifying ACs with no UI surface
ARCHITECTURE-DECISIONS.md # architecture decision records
Tech Stack
| Technology | Purpose |
|---|---|
| TypeScript | type-safe source code |
| Node.js >= 18 | runtime |
| Playwright CLI | browser control for AI agents |
| Claude Code | AI reasoning and session orchestration |
| Handlebars | HTML report templating |
| Zod | runtime schema validation |
| Commander | CLI argument parsing |
| Vitest | unit and integration testing |
| tsup | TypeScript bundling |
| YAML | configuration and knowledge base format |
Key Principles
- Systematic exploration, not random clicking
- Understand the business first, then test what matters
- Risk-rank features: P0 (40% time), P1 (30%), P2 (20%), P3 (10%)
- Apply heuristics as thinking tools, not checklists
- One bug = one report, with mandatory business impact
- Log findings in real-time with WHY reasoning, not just WHAT
- Cap sessions at 45 min to avoid context overflow
- Save phase findings to disk between phases
- Verify data integrity after every state-changing action
- Never store credentials in YAML -- use env vars via .env
License
MIT -- see LICENSE for details.
Reviews (0)
Sign in to leave a review.
Leave a reviewNo results found