Turn any web page into a clean, organized note in your Obsidian vault — in one command.
Point it at a company team page or a professor's faculty profile, and it pulls out what matters, tidies it into Markdown with proper tags and links, and files it in the right folder of your vault. No copy‑paste, no messy formatting, no dead ends when a site is built with heavy JavaScript.
It's open source, runs entirely on your machine, and is friendly to the sites it visits.
- 📥 Straight into Obsidian. Every result is a real
.mdnote with YAML frontmatter, tags, and[[wikilinks]]— it shows up in your graph immediately. - 🗂️ Sorts itself. People go in
Contacts/, profiles inProfiles/, and research inResearch/. You don't organize anything. - 👤 Knows people. Scrape a professor once and get two linked notes: their contact details and a list of their publications, each with a link back to where it was found.
- 🕸️ Crawls, not just clicks. Point it at a hub page — a faculty directory, a team page, a product listing, a docs root — and it follows the links on that page and scrapes each one, building a parent "map" note that links to every child it found.
- 🌐 Handles modern sites. Static pages are fetched fast; JavaScript‑heavy pages are rendered in a real browser automatically when needed.
- 🗺️ A living index. An
Index.mdmap-of-content links every note by category and refreshes itself after every scrape — never stale, never hand-maintained. - 🔗 A graph that connects. A person's publications link to their contact note, and colleagues share a company/institution hub — so Obsidian's graph view is actually meaningful, not a scatter of islands.
- 🤝 A polite guest. Obeys each site's
robots.txt, rotates browser identities, spaces out requests per site, and retries gently — so it behaves like a considerate visitor, not a hammer. - 🔒 Yours, locally. Your notes stay in your vault on your computer. Nothing is uploaded anywhere.
You'll need Python 3.11+ and uv (a fast Python package manager).
uv sync --extra dev # install everythingThen scrape something:
# Grab someone's contact details
uv run scraper scrape "https://example.com/team/jane" --task contact
# Peek at the note first without saving it
uv run scraper scrape "https://example.com/team/jane" --task contact --no-writeYour notes appear inside the bundled Obsidian vault at webscraper/ — already set
up with Contacts/, Profiles/, and Research/ folders, a
Welcome note, and a live Index. Open that folder as a vault in Obsidian
and everything is there. Want to use your own vault instead? Add
--vault /path/to/YourVault to any command, or set vault_path in config.yaml
(the scraper creates the category folders automatically on first write).
This is the standout feature. Say you're researching Professor X:
uv run scraper person "https://university.edu/faculty/professor-x"You get two notes, automatically linked together:
Contacts/Professor X.md— email, department/company, phone, social linksResearch/Professor X - Research.md— every publication it can find, each with a link to its source
Open either one in Obsidian and the [[Professor X]] link ties them together in
your graph. Build a research library one person at a time.
Have a page full of links, or a file of URLs? Do them all in one go:
# Several URLs
uv run scraper batch "https://a.com" "https://b.com" --task contact
# Or from a file (one URL per line)
uv run scraper batch --file urls.txt --task contactIt runs them in parallel (politely), writes a note for each success, and tells you which ones didn't pan out — one bad link never stops the rest.
Handy flags on scrape, person, and batch: --overwrite (replace an existing
note instead of making a numbered copy) and --index (refresh the vault index
right after writing).
See everything it can do:
uv run scraper tasks # list the scrape types: contact, profile, research
uv run scraper --help # full command referenceThe commands above scrape pages you already have links to. crawl goes one step
further: give it a single hub page — a department faculty list, a company team
page, a product directory, a lecture outline, a documentation root — and it finds
the relevant links on that page and scrapes each one for you.
# Scrape a faculty directory and every professor it links to
uv run scraper crawl "https://university.edu/department/faculty"
# Go two levels deep, and force every child to be a contact note
uv run scraper crawl "https://acme.com/team" --depth 2 --task contactYou end up with:
- A hub note (in
Hubs/) that acts as a table of contents, linking to every page it scraped. - One note per child page, sorted into the usual folders, each linking back to the hub — so in Obsidian's graph the whole set shows up as one tidy cluster instead of scattered notes.
It's smart about what it follows: it grabs the real links in the page's content
and skips menus, footers, share buttons, and links off to other websites. By
default it detects the best note type for each page automatically (--task auto),
but you can force one with --task. A couple of guardrails keep runs sane:
--depth 1 # how many link-levels to follow (default: 1)
--max-pages 20 # never scrape more than this many pages (default: 20)
--allow-external # opt in to following links to other domainsAnd of course it stays polite the whole time — same robots.txt respect, per-site
rate limiting, and gentle retries as everything else, and it never scrapes the
same page twice in a run.
Every scrape automatically refreshes Index.md at the top of your vault — a
section per category (Hubs, Contacts, Profiles, Research) listing every
note as a clickable [[link]]. Open it in Obsidian and everything is one hop away;
you never have to maintain it.
Prefer to do it by hand? Rebuild it anytime with uv run scraper index, or turn
the automatic refresh off per‑run with --no-index (or set auto_index: false
in config.yaml).
Notes are wired together on purpose, so Obsidian's graph view is genuinely useful:
- Scrape a person and their research note links straight to their contact note, so their details and their publications sit together in the graph.
- Everyone at the same company or institution links to a shared hub note, so colleagues cluster together.
Some sites render their content in the browser with JavaScript. If a normal fetch comes back empty, the scraper can render the page in a real headless browser and try again. Set that up once:
uv sync --extra dev --extra dynamic
uv run playwright install chromiumThen it just works — the scraper escalates to the browser on its own, or you can
force it with --dynamic.
Sensible defaults live in config.yaml — folder names, politeness (request
rate), and timeouts. Override any of it without editing files using
SCRAPER_‑prefixed environment variables, for example:
SCRAPER_VAULT_PATH=~/Notes/Vault # where notes are written
SCRAPER_THROTTLE__RATE=0.5 # go gentler: 1 request every 2 seconds
SCRAPER_FETCH__TIMEOUT=30 # longer request timeoutThis tool is for gathering information you're allowed to access. It honors
robots.txt by default — if a site asks crawlers to stay out of a path, the
scraper won't fetch it. Beyond that: respect each site's terms of service, don't
hammer servers, and don't collect personal data you have no right to. The
built‑in rate limiting and robots.txt checks are there to help you be a good
citizen — please keep them on. (You can disable the robots check with
SCRAPER_FETCH__RESPECT_ROBOTS=false, but think twice before you do.)
Adding a new kind of note is a small, well‑defined job — a schema, a parser, and a folder. Run the checks before opening a pull request:
uv run pytest && uv run ruff check && uv run mypy srcMIT — free to use, modify, and share.