# Tross Assignment - LinkedIn Scraper ## Update I wasn't selected for the role. Here's the full email I got -- ``` Hi Nikhil, Thank you for taking the time to apply to Tross and complete the assignment. We appreciate the effort you put into the process. We’ve now concluded this round of engineering hiring. For this role, we were looking for a specific mix of reverse engineering and software engineering experience, and have decided to move forward with a candidate whose experience aligns with our current requirements. As we continue to grow our engineering team, there may be opportunities in the future where your profile could be a strong fit. We’d be happy to keep your application in consideration for upcoming roles. Thank you again for your interest in Tross, and we wish you all the best. Best, Tross Careers Team ``` Async service that scrapes a public LinkedIn profile and returns it as a validated Pydantic `Profile`. It reads LinkedIn's RSC (React Server Components) wire format, not HTML, and rebuilds the full profile: top card plus About, Experience, Education, Skills, Certifications, Languages. Exposed through a small FastAPI app and secured by an admin API key. The full output shape is in `src/schemas.py`. ## Setup Requirements: Python 3.14 and PDM. ```bash git clone git@tangled.org:did:plc:26n46rcm5csiruar6lrzf2fe tross && cd tross pdm install cp config.example.toml config.toml # then edit it ``` ## Configuration All config lives in `config.toml` at the project root, loaded by `src/config.py`. Environment variables override the file. The secondary tross key is there becuase it's personal account's cookies and misuse might lead to a ban. ```toml admin_key = "..." # primary key, never expires tross_key = "..." # secondary key tross_key_valid_until = 2026-09-01 # rejected from Sept 2 onward [linkedin] [linkedin.cookies] li_at = "AQED..." # required JSESSIONID = '"ajax:..."' # required (keep the quotes) ``` Click on any profile while your network tab is open. You can right click the call made right after the click and then copy it as bash. Then head to curlconvertor and convert it to python's request. The cookie dict should be as toml dict. The app derives the `csrf-token` from `JSESSIONID` at startup. The app refuses to start if `admin_key` or `tross_key` is empty, or if `li_at` or `JSESSIONID` is missing. ## Running Local: ```bash uvicorn app:app ``` Docker (mounts `config.toml` read-only, secrets never enter the image): ```bash docker compose up -d --build ``` Swagger is at `/docs`, ReDoc is at `/redoc` (show redoc some love) health endpoint is at `/health` ## API ### `GET /api/v1/scrape/` Scrape a LinkedIn profile. - Query: `url` (a `https://www.linkedin.com/in//` URL) - Header: `X-Admin-Key` (a valid `admin_key` or `tross_key`) ```bash curl -H "X-Admin-Key: $ADMIN_KEY" \ "http://localhost:8000/api/v1/scrape/?url=https://www.linkedin.com/in/heli-v-48b726106/" ``` Returns `200` with a `Profile` JSON object. Errors use a structured body `{"error_code": "...", "message": "..."}`: | Status | error_code | When | |--------|---------------------|----------------------------------------| | 401 | UNAUTHORIZED | Missing, invalid, or expired key | | 400 | BAD_REQUEST | Malformed request (service layer) | | 404 | NOT_FOUND | Resource not found (service layer) | | 503 | SERVICE_UNAVAILABLE | Transient upstream failure (service) | | 500 | none | Unexpected error (LinkedIn 5xx, parse) | ### `GET /health` Unauthenticated liveness probe. Returns `{"status": "ok"}`. ## Approach LinkedIn's profile page is an RSC stream: `chunkId:JSON` lines with `$Lxx` / `$xx` references between chunks. The browser loads the full sections as separate async POSTs to `/flagship-web/rsc-action/actions/component`. The scraper copies that flow: ``` GET /flagship-web/in// -> top card (name, headline, images) POST /rsc-action/.../component x5 -> about, experience, education+certs, languages, skills (fetched concurrently) parse RSC stream -> resolve refs -> structure items -> Profile ``` Steps: parse chunks, resolve references, walk the React tree for text and images, group flat text into items by component structure, then structure each section into typed objects and validate as a `Profile`. Hard parts and fixes: - TLS fingerprinting: `httpx` is rejected on the POST endpoints (GET 200, POST 500). The scraper uses `curl_cffi` with `impersonate="firefox"` to present a real browser TLS fingerprint. - Reverse-engineered POSTs: needs `componentId` / `sduiid` / `parentSpanId` params, per-component payloads, `csrf-token = JSESSIONID` (quotes stripped), and page-instance headers copied from the top card. - Duplicate chunks: the stream re-emits the same content in several root chunks. `_dedupe_repeated` collapses exact block repetitions. Sections that repeat field values across items (certifications share issuer and issued date, languages can share proficiency) are de-triplicated, not value-deduplicated. - Flat text to items: a `componentKey` stack groups items (wrapper at depth 1, items at depth 2). Certifications are split on their terminal field line (`Credential ID`, or `Issued` when there is no credential id). ## Known limitations - Multi-role experience: LinkedIn nests roles under one company. The scraper returns one entry with the company in `title` and roles in `description`. - Self-view profiles: some sections may be empty or show "Add ..." placeholders. - Certification ordering: the issuer is the line after the name; rare reorderings can mislabel it. Certs without a `Credential ID` split on `Issued` instead. - Single image per field: `profile_image` and `background_image` return one rendition URL, not all size variants. - No batching or rate-limiting or backoff. - Cookies: if `li_at` expires, requests fail. There is no auto re-login (LinkedIn blocks it and it risks the account). - No caching: every request re-fetches from LinkedIn. - Best-effort parsing: if LinkedIn changes the RSC layout, the parser may need adjustment.