yt-comment-scraper #
Scrape YouTube comments via the Data API v3. Store raw responses verbatim, query them with DuckDB, optionally run sentiment scoring.
Setup #
# Core scraper + DuckDB views
uv sync
# With Cardiff XLM-R sentiment scoring
uv sync --extra sentiment
Put your API key in .env:
YOUTUBE_API_KEY=AIza...
Usage #
1. Configure channels #
Create channels.yaml:
channels:
- handle: "@ChannelName"
start_date: "2024-01-01"
end_date: "2025-01-01"
min_comments: 0
- handle: "UCxxxxxxxxxxxxxxxxxxxxxx" # channel ID also works
start_date: "2023-06-01"
start_date/end_date: ISO date strings. Omitend_dateto use now.min_comments: skip videos with fewer comments during the comments phase.
2. Scrape #
from yt_comment_scraper import ScraperConfig, YouTubeScraper
config = ScraperConfig.from_yaml("channels.yaml")
scraper = YouTubeScraper(config, output_dir="data")
scraper.run()
Resume after interruption — each phase checkpoints progress per channel.
3. Build DuckDB views #
from yt_comment_scraper import build_duckdb
build_duckdb("data") # creates data/youtube.duckdb
Then query:
import duckdb
con = duckdb.connect("data/youtube.duckdb")
con.execute("SELECT video_id, title, view_count FROM videos LIMIT 5").df()
con.execute("SELECT author_display_name, text FROM comments LIMIT 5").df()
4. Sentiment scoring (optional) #
from yt_comment_scraper.sentiment import score_comments
score_comments("data") # appends to data/youtube.duckdb
Incremental — skips already-scored comments. Configurable via environment variables:
| Variable | Default | What it does |
|---|---|---|
CARDIFF_BATCH_SIZE |
128 | Model inference batch size |
CARDIFF_FETCH_SIZE |
8192 | Comments loaded per DB fetch |
CARDIFF_DEVICE |
auto | auto, cpu, mps, cuda |
CARDIFF_COMMENT_LIMIT |
— | Cap new comments scored this run |
CARDIFF_FORCE_RERUN |
— | Set to 1 to delete existing scores and re-run |
Output table: comment_sentiment_cardiff with columns comment_id, model_name, prob_negative, prob_neutral, prob_positive, sentiment_score (positive − negative), sentiment_label, scored_at.
After scoring, re-run build_duckdb() to create two convenience views:
comments_cardiff_sentiment— every comment joined with its sentiment scores (NULLs for unscored)authors_cardiff_sentiment— per-author aggregation: mean probabilities, mean sentiment score, majority-vote label
Data layout #
data/
├── raw/
│ ├── manifest_*.json # per-channel checkpoint state
│ ├── videos/{video_id}.json
│ ├── comments/{video_id}.ndjson
│ └── channels/batch_*.ndjson
└── youtube.duckdb # views over raw files + sentiment table
Raw files are exact API responses. DuckDB views are rebuilt from them on every build_duckdb() call — no transformation happens at scrape time.
Phases #
| Phase | What it does | Resumable |
|---|---|---|
| discover | Enumerate video IDs from channel uploads playlist | ✓ |
| videos | Fetch full video metadata (snippet, statistics, etc.) | ✓ |
| comments | Fetch all commentThreads + replies per video | ✓ |
| channels | Batch-fetch metadata for unique commenters | ✓ |
Skip phases you don't need:
scraper.run(phases=["discover", "videos"])
Quota #
YouTube API daily quota is 10,000 units. Rough costs:
| Operation | Cost per call |
|---|---|
| channels.list | 1 |
| playlistItems.list | 1 |
| videos.list (50 IDs) | 1 |
| commentThreads.list | 1 |
| comments.list | 1 |
A channel with ~1,800 videos and ~400K comments costs ~6,000 units — fits in one day.