wesearch
Health Warn
- License — License: Apache-2.0
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 6 GitHub stars
Code Fail
- rm -rf — Recursive force deletion command in .github/workflows/package-validation.yml
- rm -rf — Recursive force deletion command in .github/workflows/sync-to-source.yml
Permissions Pass
- Permissions — No dangerous permissions requested
No AI report is available for this listing yet.
A wascally web research toolkit for agents.
wesearch🕸️
A wascally web research toolkit for agents; web search, web fetch, and publication research.
Quick Start
# Mac:
# # Required for quick install.
# brew install uv
# Ubuntu/Debian:
# # Required for quick install.
# sudo apt-get install -y curl
# curl -LsSf https://astral.sh/uv/install.sh | sh
uv add wesearch
# Alternatively: python -m pip install wesearch
wesearch is a synchronous, batteries-included library for programmatic web access: run a
search, fetch a page through a real-browser fingerprint, extract content, and look up scholarly
papers across multiple providers. It is the web layer factored out of a larger agent stack, so
it is built to survive bot-detection, rate limits, and flaky endpoints without a running browser
in the common case.
Example
from wesearch.search.search import search
from wesearch.fetch import RequestParams, Retry, fetch
from wesearch.scrape import get_element_content
# Web search (DuckDuckGo by default)
hits = search("denoising recursion models", num_results=10)
for r in hits:
print(r.title, r.url)
# Fetch + scrape
body, _session = fetch(
"https://example.com", request=RequestParams(retry=Retry(timeout_sec=10))
)
title = get_element_content(body.decode("utf-8"), "h1")
# Scholarly papers: Semantic Scholar + OpenAlex, reciprocal-rank-fused by default
from wesearch.paper.search import search as paper_search
from wesearch.paper.ids import normalize_id
from wesearch.paper.details import metadata
result = paper_search("attention is all you need", limit=5)
for rec in result.records:
print(rec.title, rec.year)
meta = metadata(*normalize_id("arXiv:1706.03762"))
Each name is imported from the submodule that defines it; the top-level __init__ re-exports
nothing.
What's inside
wesearch/
├── types/ the vocabulary every layer shares; imports nothing internal
│ ├── params.py RequestParams + Content/Retry/Observe/Policy; Transport,
│ │ Extractor, Trust
│ ├── extractor.py the Extract protocol each extractor satisfies
│ └── errors.py FetchError, BotDetectionError + subclasses
├── fetch/ the sole HTTP egress
│ ├── fetch.py fetch(url, request=RequestParams(...)) -> (body, session)
│ ├── transport/ how bytes are retrieved
│ │ ├── curl.py curl-cffi transport (TLS/JA3 browser impersonation)
│ │ ├── stdlib.py dependency-free urllib transport
│ │ ├── zendriver.py opt-in real-Chrome backend for JS-gated pages
│ │ └── transport_routing.py per-domain transport selection
│ ├── extractor/ how a fetched page becomes text
│ │ ├── html2text.py every text node as Markdown (the default)
│ │ ├── trafilatura.py the scored article body only; article-shaped pages
│ │ ├── raw.py the source, untouched
│ │ └── markdownify.py the document's elements as Markdown
│ ├── providers/ per-site fetch strategies (reddit, google_news, x, ...)
│ └── challenge.py bot-challenge detection and classification
├── search/ web search over pluggable backends
│ └── search.py search(...) over SearXNG / DuckDuckGo / Google;
│ SearchResult / PaperResult / ImageResult
├── paper/ scholarly-paper lookup
│ ├── search.py search(...) across Semantic Scholar, OpenAlex, SearXNG
│ ├── details.py metadata / references / citations
│ ├── authors.py author search and publication lists
│ ├── fetch.py PDF download cascade
│ └── providers/ per-source backends (openalex, s2, searxng)
├── mcp/ the MCP surface; the only place the mcp SDK is imported
│ └── server.py wesearch-mcp, one tool per public function
├── chrome/ real-browser fingerprints
│ ├── headers.py Chrome request headers (incl. x-browser-validation)
│ └── useragents.py vendored, refreshable User-Agent pools
├── web.py fetch_web(...): fetch + provider dispatch + extraction
├── profile.py cross-process per-(egress_ip, domain) cookie + UA jar
├── ratelimit.py cross-process, per-domain rate limiting
└── scrape.py get_element_content(html, selector)
Transports
fetch picks a transport per domain:
- curl-cffi (default): TLS/JA3 browser impersonation, no browser process.
- stdlib: dependency-free
urllibfallback. - zendriver: an opt-in headless-Chrome backend for JavaScript-gated pages, used only for
domains that require it.
A persistent per-(egress_ip, domain) profile (cookies + User-Agent) is loaded and saved
transparently, and cross-process rate limiting paces requests so concurrent workers stay under
each site's threshold.
API keys & configuration
Everything works keyless out of the box; the environment variables below raise your rate limits
or unlock a backend, and are all optional unless noted.
SEMANTIC_SCHOLAR_API_KEY-- optional. Without it, Semantic Scholar lookups (paper.search,paper.details) share a low-rate public tier, andpaper.search(..., source="fused")may
returncomplete=Falsewhen S2 throttles the request. Set it for a higher-throughput tier.OPENALEX_EMAIL-- optional. Identifies you to OpenAlex's "polite pool" for more headroom than
anonymous requests.OPENALEX_API_KEY-- optional. A higher OpenAlex request budget.SEARXNG_URL-- required only to use a"searxng"backend (search(..., backend="searxng")
orpaper.search(..., source="searxng")); the base URL of a SearXNG instance you control or
trust.
State (the cookie/User-Agent profile jar, cross-process rate-limit lockfiles, the browser
transport's persistent Chrome profile) is written under the OS's standard per-user data
directory -- XDG_DATA_HOME (or ~/.local/share) on Linux, ~/Library/Application Support on
macOS, %LOCALAPPDATA% on Windows -- namespaced per component (see wesearch/lib/userdirs.py).
No configuration file is required or read.
The "zendriver" transport needs a non-snap Chrome or Chromium (e.g. Google Chrome's.deb on x86_64). Snap-packaged Chromium fails every launch with BrowserUnavailableError
for two reasons, neither of which the library can work around:
- The snap wrapper takes several seconds to expose DevTools -- far beyond the launch budget
(0.5s per attempt, 6 attempts), which is sized so thecurl-then-zendrivercascade fails
fast on hosts with no usable browser rather than stalling every fetch. - Snap's AppArmor confinement silently blocks writes under hidden home paths like
~/.local/share, so Chrome dies on its profile lock wherever the profile jar lands by
default.
Development
uv sync --all-groups
Tests are tiered with pytest markers; the default run (uv run pytest) executes only the fast
unit tier:
addopts = -m 'not ci_smoke and not cuda and not integration and not performance and not cluster and not slow'
ci_smoke-- slower package smoke tests, run explicitly in CI.cuda-- requires a real CUDA device.integration-- requires networking or external CLIs.performance-- timing-sensitive.cluster-- requires live cluster access.slow-- expensive local correctness tests (JIT, full fixtures, git, bash, a fresh interpreter).real_llm-- spawns a live LLM CLI; skipped unlessRUN_REAL_LLM=1.
Run a specific tier with uv run pytest -m integration, or everything withuv run pytest -m ''.
The zendriver backend tests need a Chrome or Chromium binary on PATH; without one they are
skipped automatically. No other system dependency is required to run the fast unit tier.
MCP server
The toolkit is directly callable by coding agents over MCP:
pip install 'wesearch[mcp]'
wesearch-mcp # serves stdio; register it with your MCP client
For Claude Code: claude mcp add --scope user wesearch -- wesearch-mcp.
Tools exposed: paper_search (fused Semantic Scholar + OpenAlex),paper_details, paper_references, paper_citations, paper_pdf
(downloads into the user cache and returns the path), author_search,author_papers, web_search, and web_fetch (extracted page text; the
default transport escalates to a headless browser when a site bot-blocks the
header transports). Outputs are
deliberately compact for model consumption; abstracts are truncated and
empty fields dropped. The server is synchronous and per-client — state that
must be shared (rate limits, cookie/UA profiles) is already cross-process
safe on disk, so no daemon is needed.
See also
Sibling projects in the rekursiv-ai family:
- sagent — The self-mutating multi-provider coding-agent CLI and typed Python library.
- trackinizer — Centralized agent database for tracking inquiries, work, and the evidence behind conclusions.
- madcatter — Rich-based Markdown renderer for the terminal; ships the
mdcatCLI. - priml — Composable PyTorch building blocks: models, optimizers, losses, and a step-based training loop.
- configgle — Hierarchical experiment configuration in typed pure-Python dataclasses instead of YAML.
- copybarista — Bidirectional source sync for publishing OSS-ready trees from a monorepo.
- sudoku — Sudoku-Extreme solved end to end with a 7M-parameter recursive transformer.
Citing
If you find our work useful, please consider citing:
@misc{rekursivai2026wesearch,
title={Wesearch - A wascally web research toolkit for agents.},
author={Joshua V. Dillon},
year={2026},
howpublished={Github},
url={https://github.com/rekursiv-ai/wesearch},
}
Reviews (0)
Sign in to leave a review.
Leave a reviewNo results found