Skip to content

Repository files navigation

WebScraper → Obsidian 🕸️📓

Turn any web page into a clean, organized note in your Obsidian vault — in one command.

Point it at a company team page or a professor's faculty profile, and it pulls out what matters, tidies it into Markdown with proper tags and links, and files it in the right folder of your vault. No copy‑paste, no messy formatting, no dead ends when a site is built with heavy JavaScript.

It's open source, runs entirely on your machine, and is friendly to the sites it visits.


Why you'll like it

  • 📥 Straight into Obsidian. Every result is a real .md note with YAML frontmatter, tags, and [[wikilinks]] — it shows up in your graph immediately.
  • 🗂️ Sorts itself. People go in Contacts/, profiles in Profiles/, and research in Research/. You don't organize anything.
  • 👤 Knows people. Scrape a professor once and get two linked notes: their contact details and a list of their publications, each with a link back to where it was found.
  • 🕸️ Crawls, not just clicks. Point it at a hub page — a faculty directory, a team page, a product listing, a docs root — and it follows the links on that page and scrapes each one, building a parent "map" note that links to every child it found.
  • 🌐 Handles modern sites. Static pages are fetched fast; JavaScript‑heavy pages are rendered in a real browser automatically when needed.
  • 🗺️ A living index. An Index.md map-of-content links every note by category and refreshes itself after every scrape — never stale, never hand-maintained.
  • 🔗 A graph that connects. A person's publications link to their contact note, and colleagues share a company/institution hub — so Obsidian's graph view is actually meaningful, not a scatter of islands.
  • 🤝 A polite guest. Obeys each site's robots.txt, rotates browser identities, spaces out requests per site, and retries gently — so it behaves like a considerate visitor, not a hammer.
  • 🔒 Yours, locally. Your notes stay in your vault on your computer. Nothing is uploaded anywhere.

Quick start

You'll need Python 3.11+ and uv (a fast Python package manager).

uv sync --extra dev            # install everything

Then scrape something:

# Grab someone's contact details
uv run scraper scrape "https://example.com/team/jane" --task contact

# Peek at the note first without saving it
uv run scraper scrape "https://example.com/team/jane" --task contact --no-write

Your notes appear inside the bundled Obsidian vault at webscraper/ — already set up with Contacts/, Profiles/, and Research/ folders, a Welcome note, and a live Index. Open that folder as a vault in Obsidian and everything is there. Want to use your own vault instead? Add --vault /path/to/YourVault to any command, or set vault_path in config.yaml (the scraper creates the category folders automatically on first write).


The people workflow ✨

This is the standout feature. Say you're researching Professor X:

uv run scraper person "https://university.edu/faculty/professor-x"

You get two notes, automatically linked together:

  • Contacts/Professor X.md — email, department/company, phone, social links
  • Research/Professor X - Research.md — every publication it can find, each with a link to its source

Open either one in Obsidian and the [[Professor X]] link ties them together in your graph. Build a research library one person at a time.


Scrape a whole list at once

Have a page full of links, or a file of URLs? Do them all in one go:

# Several URLs
uv run scraper batch "https://a.com" "https://b.com" --task contact

# Or from a file (one URL per line)
uv run scraper batch --file urls.txt --task contact

It runs them in parallel (politely), writes a note for each success, and tells you which ones didn't pan out — one bad link never stops the rest.

Handy flags on scrape, person, and batch: --overwrite (replace an existing note instead of making a numbered copy) and --index (refresh the vault index right after writing).

See everything it can do:

uv run scraper tasks       # list the scrape types: contact, profile, research
uv run scraper --help      # full command reference

Crawl a whole hub in one go 🕸️

The commands above scrape pages you already have links to. crawl goes one step further: give it a single hub page — a department faculty list, a company team page, a product directory, a lecture outline, a documentation root — and it finds the relevant links on that page and scrapes each one for you.

# Scrape a faculty directory and every professor it links to
uv run scraper crawl "https://university.edu/department/faculty"

# Go two levels deep, and force every child to be a contact note
uv run scraper crawl "https://acme.com/team" --depth 2 --task contact

You end up with:

  • A hub note (in Hubs/) that acts as a table of contents, linking to every page it scraped.
  • One note per child page, sorted into the usual folders, each linking back to the hub — so in Obsidian's graph the whole set shows up as one tidy cluster instead of scattered notes.

It's smart about what it follows: it grabs the real links in the page's content and skips menus, footers, share buttons, and links off to other websites. By default it detects the best note type for each page automatically (--task auto), but you can force one with --task. A couple of guardrails keep runs sane:

--depth 1            # how many link-levels to follow (default: 1)
--max-pages 20       # never scrape more than this many pages (default: 20)
--allow-external     # opt in to following links to other domains

And of course it stays polite the whole time — same robots.txt respect, per-site rate limiting, and gentle retries as everything else, and it never scrapes the same page twice in a run.


One map for your whole vault 🗺️

Every scrape automatically refreshes Index.md at the top of your vault — a section per category (Hubs, Contacts, Profiles, Research) listing every note as a clickable [[link]]. Open it in Obsidian and everything is one hop away; you never have to maintain it.

Prefer to do it by hand? Rebuild it anytime with uv run scraper index, or turn the automatic refresh off per‑run with --no-index (or set auto_index: false in config.yaml).

A graph that actually connects

Notes are wired together on purpose, so Obsidian's graph view is genuinely useful:

  • Scrape a person and their research note links straight to their contact note, so their details and their publications sit together in the graph.
  • Everyone at the same company or institution links to a shared hub note, so colleagues cluster together.

Handling JavaScript‑heavy sites

Some sites render their content in the browser with JavaScript. If a normal fetch comes back empty, the scraper can render the page in a real headless browser and try again. Set that up once:

uv sync --extra dev --extra dynamic
uv run playwright install chromium

Then it just works — the scraper escalates to the browser on its own, or you can force it with --dynamic.


Configuration

Sensible defaults live in config.yaml — folder names, politeness (request rate), and timeouts. Override any of it without editing files using SCRAPER_‑prefixed environment variables, for example:

SCRAPER_VAULT_PATH=~/Notes/Vault      # where notes are written
SCRAPER_THROTTLE__RATE=0.5            # go gentler: 1 request every 2 seconds
SCRAPER_FETCH__TIMEOUT=30             # longer request timeout

Please scrape responsibly

This tool is for gathering information you're allowed to access. It honors robots.txt by default — if a site asks crawlers to stay out of a path, the scraper won't fetch it. Beyond that: respect each site's terms of service, don't hammer servers, and don't collect personal data you have no right to. The built‑in rate limiting and robots.txt checks are there to help you be a good citizen — please keep them on. (You can disable the robots check with SCRAPER_FETCH__RESPECT_ROBOTS=false, but think twice before you do.)


Contributing

Adding a new kind of note is a small, well‑defined job — a schema, a parser, and a folder. Run the checks before opening a pull request:

uv run pytest && uv run ruff check && uv run mypy src

License

MIT — free to use, modify, and share.

About

Production grade web scraper.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages