[READ-ONLY] Mirror of https://github.com/agbocsardi/yt-comment-scraper.
Python 100%

README.md

yt-comment-scraper #

Scrape YouTube comments via the Data API v3. Store raw responses verbatim, query them with DuckDB, optionally run sentiment scoring.

Setup #

# Core scraper + DuckDB views
uv sync

# With Cardiff XLM-R sentiment scoring
uv sync --extra sentiment

Put your API key in .env:

YOUTUBE_API_KEY=AIza...

Usage #

1. Configure channels #

Create channels.yaml:

channels:
  - handle: "@ChannelName"
    start_date: "2024-01-01"
    end_date: "2025-01-01"
    min_comments: 0

  - handle: "UCxxxxxxxxxxxxxxxxxxxxxx"  # channel ID also works
    start_date: "2023-06-01"
  • start_date / end_date: ISO date strings. Omit end_date to use now.
  • min_comments: skip videos with fewer comments during the comments phase.

2. Scrape #

from yt_comment_scraper import ScraperConfig, YouTubeScraper

config = ScraperConfig.from_yaml("channels.yaml")
scraper = YouTubeScraper(config, output_dir="data")
scraper.run()

Resume after interruption — each phase checkpoints progress per channel.

3. Build DuckDB views #

from yt_comment_scraper import build_duckdb

build_duckdb("data")  # creates data/youtube.duckdb

Then query:

import duckdb
con = duckdb.connect("data/youtube.duckdb")
con.execute("SELECT video_id, title, view_count FROM videos LIMIT 5").df()
con.execute("SELECT author_display_name, text FROM comments LIMIT 5").df()

4. Sentiment scoring (optional) #

from yt_comment_scraper.sentiment import score_comments

score_comments("data")  # appends to data/youtube.duckdb

Incremental — skips already-scored comments. Configurable via environment variables:

Variable Default What it does
CARDIFF_BATCH_SIZE 128 Model inference batch size
CARDIFF_FETCH_SIZE 8192 Comments loaded per DB fetch
CARDIFF_DEVICE auto auto, cpu, mps, cuda
CARDIFF_COMMENT_LIMIT — Cap new comments scored this run
CARDIFF_FORCE_RERUN — Set to 1 to delete existing scores and re-run

Output table: comment_sentiment_cardiff with columns comment_id, model_name, prob_negative, prob_neutral, prob_positive, sentiment_score (positive − negative), sentiment_label, scored_at.

After scoring, re-run build_duckdb() to create two convenience views:

  • comments_cardiff_sentiment — every comment joined with its sentiment scores (NULLs for unscored)
  • authors_cardiff_sentiment — per-author aggregation: mean probabilities, mean sentiment score, majority-vote label

Data layout #

data/
├── raw/
│   ├── manifest_*.json        # per-channel checkpoint state
│   ├── videos/{video_id}.json
│   ├── comments/{video_id}.ndjson
│   └── channels/batch_*.ndjson
└── youtube.duckdb             # views over raw files + sentiment table

Raw files are exact API responses. DuckDB views are rebuilt from them on every build_duckdb() call — no transformation happens at scrape time.

Phases #

Phase What it does Resumable
discover Enumerate video IDs from channel uploads playlist ✓
videos Fetch full video metadata (snippet, statistics, etc.) ✓
comments Fetch all commentThreads + replies per video ✓
channels Batch-fetch metadata for unique commenters ✓

Skip phases you don't need:

scraper.run(phases=["discover", "videos"])

Quota #

YouTube API daily quota is 10,000 units. Rough costs:

Operation Cost per call
channels.list 1
playlistItems.list 1
videos.list (50 IDs) 1
commentThreads.list 1
comments.list 1

A channel with ~1,800 videos and ~400K comments costs ~6,000 units — fits in one day.