sharpen

agent
Security Audit
Warn
Health Warn
  • License — License: Apache-2.0
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 5 GitHub stars
Code Pass
  • Code scan — Scanned 12 files during light audit, no dangerous patterns found
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

An agentic-AI, full-featured quant strategy builder — from a concept to a validated, trained and deployment-ready system: automated alpha mining, backtesting and falsification, and reinforcement-learning trading agents.

README.md

Sharpen

Support on Ko-fi

Sharpen is an agentic-AI, full-featured quant strategy builder — from a concept to a validated, trained and deployment-ready system.

Describe what you want in any natural language. Sharpen turns it into a pre-registered spec,
writes the signal and the tests, trains it, and then tries to break it — with every result
deflated for the number of things you tried.

Sharpen is built to be driven by an AI coding agent such as Claude Code or Codex. The rules the
agent follows ship in the repository (CLAUDE.md): how to
turn an idea into a pre-registered spec, which gates it must pass, how RL training is staged, and
what must never happen (look-ahead, cross-split normalization, fused train-and-evaluate runs). You
say what you want, in whatever language you work in; the agent writes the signal, the config and
the tests, runs the pipeline, and stops at every gate that fails.

Underneath is a full research-to-execution stack:

  • A seven-tier validation funnel that deflates every result for the number of things you tried.
  • A full backtesting and falsification toolkit: walk-forward and recent out-of-sample
    backtests, realistic costs, timing nulls, planted-signal power tests, stress and noise
    robustness. Its job is to try to break a strategy before the market does.
  • Crucible, an automated alpha mining platform in which an LLM proposes hypotheses it can never
    see the scores of.
  • A frontier deep-RL stack: distributional SAC, CrossQ, IQN, BDQ and PPO, multi-asset and
    execution environments, distributed GPU hyper-parameter search, and a staged training protocol.
  • Paper and live execution across six broker adapters, with monitoring, drift detection and a
    kill switch.

Sharpen architecture: a coding agent under the agent contract drives six stages (pre-register, mine, validate, train, break, run), each ending at a gate, over guardrails that apply to every stage

Not investment advice. Not a trading product. Every performance figure in this repository
comes from a historical simulation or a paper run. Read DISCLAIMER.md.

Ep01: Meet Sharpen (full episode). Watch on YouTube.

Ep01: Meet Sharpen (full episode) · Watch on YouTube


From a concept to a strategy

Here is the end-to-end path, as you would drive it from a coding agent. Each step names what you
ask, what the agent does, and the gate that decides whether you go on. The requests are shown in
English, but you can write them in any language.

1. State the idea.

"Test whether 12-month time-series momentum on liquid ETFs across equities, bonds, commodities
and currencies survives costs."

The agent writes a pre-registration under docs/research/: the claim, universe, horizon, cost
model and the result that would kill it. It commits this before running anything, and the
signal spec is content-hashed so it can't be quietly edited after the results are in.

2. Build and validate the signal.

"Build it as a signal and run it through the validation funnel."

The agent implements a compute(panel) signal and scores it through the funnel:

python scripts/research/eval_signals.py --batch <batch> --panel <panel> --out results/<run>

You get a scorecard with information coefficient, net Sharpe at each cost level, deflated Sharpe,
false-discovery q-value and purged cross-validation, plus a verdict. Only PROMISING goes on.

3. Try to break it.

"Falsify it: does the rule know when to trade, or would random timing do as well? Could this
test even detect an edge of this size?"

The agent runs the falsification battery: a circular-shift timing null, a power check that
measures the smallest edge the test can detect, and cost and stress sweeps. A strategy that loses
to randomly re-timed copies of its own trades is closed here, not in production.

4. Let the machine mine alphas for you (optional).

"Run Crucible on the synthetic substrate for four nights."

python scripts/research/crucible_orchestrator.py --mode synthetic --nights 4 --force

Crucible runs continuous alpha mining on its own: it proposes, mines, deflates and
forward-incubates candidates, and forward-tests survivors only on bars that postdate the
hypothesis.

5. Build a portfolio.

"Combine the validated sleeves into a volatility-targeted book and check whether it's
overfit."

The book is assembled and put through deflated Sharpe and probability-of-backtest-overfitting
checks (scripts/research/audit_tailwind_book.py is a worked example).

6. Train an RL strategy on top.

"Train a SAC overlay that has to beat the linear book out of sample."

The agent writes a config from a reference YAML and runs SharpOps, one stage and one tracked run
at a time. Each stage writes a manifest the next stage checks:

python scripts/validate_config.py  --config configs/<cfg>.yaml --stage hpo
python scripts/run_full_pipeline.py --config configs/<cfg>.yaml --stage hpo --agent sac
# then: l1-multiseed -> ensemble-confirm -> wf -> oos -> paper-deploy

7. Deploy to paper, then watch it.

"Deploy the ensemble to the paper account and alert me if it drifts."

The paper-portfolio executor, parity harness and Docker stack (Prometheus, Grafana, Telegram)
take over. Drift detection and a kill file guard the live path.

The step-by-step guide, with every gate spelled out, is
docs/guides/building-a-strategy.md.


Features

Agent-driven workflow

  • Rules the agent actually follows. CLAUDE.md encodes the invariants, the
    training protocol, the config schema and the audit chain. Every rule in it was written after
    something went wrong.
  • Deterministic gates, not self-review. Config validation, the OHLCV cleaner, pytest, ruff and
    cross-checked profit factors are tool executions the agent must pass. It can't mark its own work
    as done.
  • A deep lifecycle audit workflow (.claude/workflows/deep_strategy_audit.js) runs a finder
    agent and a skeptic agent per pillar before anything is promoted.

Signal validation funnel — sharpen/signals/

Signals are pre-registered, content-hashed specs. Each tier catches a different way a backtest
lies:

Tier Test
0 Causality (truncation tripwire), OHLC hygiene, coverage
1 Multi-horizon IC, decile spread, breadth, decay half-life
2 Capturability: net Sharpe per cost model, cost wall, turnover
3 Subperiod stability and recent out-of-sample
3.5 Combinatorial purged path resampling with embargo
4 Deflated Sharpe, effective N, FDR/BHY, Harvey-Liu-Zhu hurdle
5 Orthogonality to a factor book (residual Sharpe)

Signal libraries: WorldQuant 101, TradingView indicators, a demo set, and a genetic DSL search
(sharpen/signals/generation/) with cohort-level Monte-Carlo null gating.

Crucible: automated alpha mining — sharpen/crucible/

Systematic alpha mining with validation inside the loop.

Crucible architecture: acquire, hypothesize, mine, deflate, incubate in a lockbox, hand off to a human audit; the proposer reads only a score-free view of the trial ledger, and every run pins a manifest of version, gates, data and seeds

Ep02: Crucible (full episode). Watch on YouTube.

Ep02: Crucible (full episode) · Watch on YouTube

  • LLM hypothesis proposer (agentic/llm_proposer.py, uses the Claude API). It turns
    natural-language priors into pre-registered specs. It is structurally blind to scores: the
    module can't reach any verdict, so the search can't overfit to its own results.
  • Free-data connectors: FRED, CFTC COT, SEC EDGAR, GDELT, Stooq, TWSE, TAIFEX.
  • Mine, deflate, incubate: every candidate is deflated against everything the search has ever
    tried, then forward-incubated in a lockbox.
  • Reproducible: gate definitions are hash-frozen, and every run records its provenance.
  • Community ledger (community/ledger/): if you mine a data set
    and find nothing, one command exports the result as counts and hashes, with no formulas or
    scores, so the next person does not repeat the search. Each entry says whether the search was
    finished (CLOSED_DECISIVE) or the test was too weak to tell (OPEN_UNDERPOWERED).
  • Power study (studies/crucible-power/): why the miner promoted zero alphas, with a 15-second rerun.

Backtesting and falsification

A backtest shows what a strategy did. These tools test whether it would do it again.

Tool What it answers
Walk-forward + recent OOS (SharpOps wf, oos stages) Does the edge hold across rolling folds and on the newest unseen data, with bootstrap confidence intervals?
Realistic costs (sharpen/signals/costs.py) Standard and harsh per-market cost models, and the cost level at which the edge disappears
Circular-shift timing null (scripts/research/hma_cross_falsification.py) Keeps exposure, trade count and holding periods, shifts the timing. Does the rule know when to trade?
Edge-sign flip + matched exposure (scripts/research/risk_overlay_lab.py) Is a risk overlay adding skill, or just de-levering?
Planted-signal power (planted_sweep.py, forward_power.py) What is the smallest edge this test can detect with this much data?
Matched-null search (null_grid_sim.py, crucible_matched_null.py) What does the same search find on data with no signal at all?
Gate reachability (execution_overlay_action_ceiling.py) Could any policy in this action space clear the gate? Answer it before spending GPU hours
Multiplicity accounting (sharpen/signals/multiplicity.py) Deflates against every hypothesis tested, however the batches were split
Deflated Sharpe + PBO (audit_tailwind_book.py) Is a portfolio's Sharpe explained by how many variants were tried?
Stress and robustness (sharpen/eval/) Fixed-lot drawdown stress, price-path noise, and config-sensitivity sweeps on a frozen policy
Profit-factor cross-check (agent rule PF-XCHECK) The agent recomputes PF from both mid price and close, and stops if they diverge by more than 30%

Frontier deep RL — sharpen/agents/, sharpen/envs/

Agents
SAC Continuous control with CrossQ-style Batch Renormalization critics, an ensemble wrapper for live inference, regime-balanced replay
Distributional SAC Quantile-regression critics for risk-aware sizing
IQN Implicit quantile networks: learns the full return distribution
BDQ Branching dueling Q-networks for multi-dimensional discrete actions
PPO Continuous and discrete variants
Environments
Continuous swing Single-asset directional trading with differential-Sharpe reward and a deadband
Multi-asset allocator Cross-asset book; the agent modulates a vol-targeted linear core it must beat out of sample
Execution scheduler RL that shapes the trade path to cut implementation shortfall
Market making Spread, skew and intensity control on limit-order-book data
Wrappers Signal-gated trading, risk shaping and prop-firm constraints, observation guards

Training engineering: torch.compile, mixed precision, mega-batch updates with tunable
update-to-data ratio, and Tensor-Core-aligned networks. Distributed Optuna HPO runs across GPU
fleets (scripts/distributed_hpo_coordinator.py, distributed_hpo_worker.py), with
bare-metal and Vast.ai deploy scripts.

SharpOps stages every RL project as
data-prep → hpo → l1-multiseed → ensemble-confirm → wf → oos → paper-deploy. Each stage is one
tracked WandB run that produces one decision artifact, and validate_config.py rejects fused
pipelines and leaky configs before a GPU spins up. See docs/sharpops.md.

Execution and operations — sharpen/paper/, sharpen/live/, docker/live/

  • Six broker adapters: Bybit perpetuals, ccxt exchanges, DXtrade, Interactive Brokers futures,
    cTrader, OANDA.
  • Paper-portfolio executor with a sim-to-live parity harness.
  • Docker stack: Prometheus, Grafana, Telegram alerting, drift detection, safe mode and a kill
    file.

Guardrails built in

Each rule is guarded by a negative test: one that fails if the defect is reintroduced.

ID Rule
LEAK-1 Normalization statistics reset at every train/val/test boundary
LEAK-2 No input at bar t carries data stamped after t, including coarse-timeframe bars and anything a gate reads
BUG-03 Hindsight-shaped reward terms are zero in backtests
CRU-1 Crucible gates are hash-frozen; a new capability can't change a past verdict
CRU-2 The hypothesis agent reads only dedup keys and killed families, never verdicts

Gate thresholds live in configs/*.gates.yaml, never in code.


Try it in 30 seconds

No data and no API keys. Python 3.11+.

pip install -e ".[dev]"
python scripts/research/eval_signals.py --batch demo --panel "synthetic:1400,60" --out results/demo

Or open the repo in your coding agent and ask for it: "Run the demo signals on a synthetic panel
and explain the scorecard."

This scores four demo signals on a synthetic panel through the full funnel (abridged columns):

| # | signal  | verdict | IC-IR  | DSR   | FDR-q | cpcvOOS | fricSh | netSh@std |
|---|---------|---------|--------|-------|-------|---------|--------|-----------|
| 1 | mom_60d | LOGGED  |  0.045 | 0.485 | 0.835 |    0.22 |   0.23 |     -0.59 |
| 2 | mom_20d | LOGGED  | -0.014 | 0.097 | 0.835 |   -0.15 |  -0.31 |     -1.66 |
| 3 | rev_5d  | LOGGED  | -0.043 | 0.018 | 0.835 |   -0.08 |  -0.27 |     -3.10 |
| 4 | vol_20d | LOGGED  | -0.054 | 0.016 | 0.835 |   -0.33 |  -0.22 |     -1.60 |

LOGGED means fully measured and below the promotion bar. That's where almost every candidate
lands, which is what makes a PROMISING worth looking at.
Getting started explains every column and walks through a full
Crucible discovery tick.


Battle-tested on its own research

Sharpen was built by running it hard: 77 strategies and probes across eight families went through
this machinery, from directional RL and options premia to market making and cross-sectional
equities. The full record, with the number that decided each one, is in
NEGATIVE_RESULTS.md.

  • It found real signal. Cross-asset time-series momentum on free daily ETF data came through
    at net Sharpe 0.60 (0.39 on 32 ETFs never used in development).
  • Every headline number reproduces from a fresh clone in about a minute on free data:
python scripts/research/xsec_momentum_falsification.py   # net Sharpe 0.601, 4/4 classes
python scripts/research/tsmom_excess_return_check.py     # 0.511 in excess of T-bills
python scripts/research/audit_tailwind_book.py           # DSR 0.896 < 0.95, PBO 0.0009
python scripts/research/value_falsification.py           # value factor -0.364, NO-GO

Expected output: docs/REPRODUCE.md.


Documentation

The hub is docs/README.md.

Document Covers
Getting started Install and two verified first runs
Building a strategy Idea → build → validate → portfolio → paper, with the gate at each step
Signal research Writing a signal; the validation funnel; the DSL
Crucible Automated alpha mining
RL pipeline SharpOps in practice
Configuration Config and gate schemas
Live trading Brokers, Docker, observability, kill switch
Data · sources and licensing Read before using real prices. A fresh clone contains no market data
Methodology Pre-registration, leakage rules, deflation, power, audits
Architecture · Testing · Troubleshooting Package map, test suite, known failure modes
Research archive ~100 pre-registrations, audits and verdicts

Project layout

sharpen/
├── signals/    Signal specs, alpha DSL, validation funnel
├── crucible/   Alpha mining: automated hypothesis search with an LLM proposer
├── agents/     SAC, distributional SAC, IQN, BDQ, PPO
├── envs/       Trading, allocator, execution and market-making environments
├── hpo/        Optuna objectives and samplers
├── training/   Trainers
├── data/       Loaders, feature engineering, splitter
├── paper/      Paper-portfolio executor and parity harness
├── crypto/  cfd/  futures/  live/   Broker adapters and live engine
└── ...         eval, monitoring, analytics, portfolio
configs/  scripts/  tests/  docs/  docker/live/
community/ledger/   Crucible searches shared by users: what was searched and whether it is closed

Session tags such as S553 in the docs refer to the private R&D log, which is not published.


Tests

python -m pytest

About 3,300 tests run by default and pass on every push in CI (Windows, CPU-only). Exact counts
vary by machine, because some tests skip when local data caches, optional dependencies or full
git history are missing. The 28
deselected tests are slow or integration tests excluded by default; at least one can hang rather
than fail, so run them individually with a timeout.


Contributing, security, and a personal note

Challenges to a result are the most useful contribution, and a Crucible search that found nothing
is worth sharing too: see CONTRIBUTING.md.
Report security issues privately per SECURITY.md. Participation follows the
Code of Conduct.
Questions and ideas go in GitHub Discussions.

EPILOGUE.md is the author's personal conclusion. It is opinion, labelled as such, and
goes beyond what this repository shows.


Support

If Sharpen is useful to you, you can support its development on
Ko-fi.
Support pays for compute and upkeep of the open research. It buys no access to strategies,
signals or advice, and nothing here is a claim about future returns.

Commercial support. Firms using Sharpen or Crucible can engage the author for integration,
custom research pipelines, independent falsification audits of a strategy, or training on the
methodology. Contact the maintainer through github.com/bigcan.
Engagements cover software and research methodology only, never investment advice, signals or
capital management.


License

Apache License 2.0. Copyright 2026 Keng Lee. Attribution for derived third-party code is
in NOTICE. No market data is distributed; see docs/DATA.md.
"Sharpen" and "Crucible" are trademarks of Keng Lee; the license covers the code, not the names.
See TRADEMARKS.md.

Reviews (0)

No results found