Skill

yc-jobs-scraper

Scrapes daily YC job listings from workatastartup.com without duplicates using Playwright auth, Inertia.js JSON extraction, HTML fallback, and SQLite deduplication. Useful for updating startup job databases.

Javascript

Node

SQLite

automation

database

npx claudepluginhub varnan-tech/opendirectory --plugin opendirectory-gtm-skills

Tool Access

This skill uses the workspace's default tool permissions.

Preview

This skill provides a robust architecture for scraping jobs from YCombinator and `workatastartup.com`. It is designed to run automatically, bypass login bottlenecks, and maintain state to never scrape duplicate jobs.

Supporting Assets

scripts/auth.jsscripts/db.jsscripts/export_radar_candidates.jsscripts/package-lock.jsonscripts/package.jsonscripts/scraper.js

SKILL.md

Similar Skills

intelligent-web-scraper

Automatically scrapes websites by analyzing page structure, handling pagination/anti-blocking, discovering article series using Playwright and Crawl4AI. Zero config needed.

15 files17 tools

intelligent-web-scraper

scraper-builder

Builds production-ready web scrapers for any website using Bright Data APIs including Web Unlocker, Browser, and SERP. Guides site analysis, selector extraction, pagination handling, and code implementation in Python or Node.js.

5 files

brightdata-plugin

data-scraper-agent

173.8k

Build a fully automated AI-powered data collection agent for any public source — job boards, prices, news, GitHub, sports, anything. Scrapes on a schedule, enriches data with a free LLM (Gemini Flash), stores results in Notion/Sheets/Supabase, and learns from user feedback. Runs 100% free on GitHub Actions. Use when the user wants to monitor, collect, or track any public data automatically.

everything-claude-code

Stats

Stars149

Forks15

Last CommitApr 14, 2026

Actions

View Source View Plugin View on GitHub View README

Help us improve

Share bugs, ideas, or general feedback.

YC Jobs Scraper

This skill provides a robust architecture for scraping jobs from YCombinator and workatastartup.com. It is designed to run automatically, bypass login bottlenecks, and maintain state to never scrape duplicate jobs.

Architecture

The scraper uses a hybrid approach to maximize reliability and minimize bot detection:

Authentication: scripts/auth.js uses Playwright to let a human log in once and saves the session to scripts/state.json.
Database: scripts/db.js uses better-sqlite3 to manage scripts/jobs.db. It tracks every company_slug and job_id ever seen.
Primary Extraction: scripts/scraper.js loads state.json, visits YC query URLs, and extracts company slugs from the hidden Inertia.js data-page JSON payload.
Job Extraction (JSON): It then visits the authenticated company pages (/companies/[slug]) to extract jobs from the backend JSON payload to ensure we get the real job_id for accurate deduplication.
Job Extraction (Fallback): If the JSON extraction fails, it falls back to parsing public HTML job cards from ycombinator.com/companies/[slug]/jobs.

Workflows

1. First-Time Setup

If this is the first time running the scraper in an environment, or if node_modules is missing:

cd @path/scripts
npm install
npx playwright install

2. Authentication (Manual Step)

If scripts/state.json is missing or expired, the scraper will fail. You must instruct the human user to run the authentication script manually:

cd @path/scripts
node auth.js

Tell the user a browser will open, and they must log in. Playwright will automatically save the cookies/tokens to state.json.

3. Running the Daily Scraper

To scrape for new companies and jobs:

cd @path/scripts
node scraper.js

This script will output exactly how many new companies and new jobs were found. Because of jobs.db, running it multiple times consecutively will result in 0 new jobs found.

4. Querying the Database

If you need to analyze the scraped data or view the companies/jobs, you can query scripts/jobs.db directly using better-sqlite3.

Example: Count Companies

cd @path/scripts
node -e "const db = require('better-sqlite3')('jobs.db'); console.log('Companies:', db.prepare('SELECT COUNT(*) as count FROM companies').get().count);"

Example: View Recent Jobs

cd @path/scripts
node -e "const db = require('better-sqlite3')('jobs.db'); const jobs = db.prepare('SELECT title, company_slug, location FROM jobs ORDER BY created_at DESC LIMIT 5').all(); console.table(jobs);"