Tross Assignment - LinkedIn Scraper #
Update #
I wasn't selected for the role. Here's the full email I got --
Hi Nikhil,
Thank you for taking the time to apply to Tross and complete the assignment. We appreciate the effort you put into the process.
We’ve now concluded this round of engineering hiring. For this role, we were looking for a specific mix of reverse engineering and software engineering experience, and have decided to move forward with a candidate whose experience aligns with our current requirements.
As we continue to grow our engineering team, there may be opportunities in the future where your profile could be a strong fit. We’d be happy to keep your application in consideration for upcoming roles.
Thank you again for your interest in Tross, and we wish you all the best.
Best,
Tross Careers Team
Async service that scrapes a public LinkedIn profile and returns it as a
validated Pydantic Profile. It reads LinkedIn's RSC (React Server Components)
wire format, not HTML, and rebuilds the full profile: top card plus About,
Experience, Education, Skills, Certifications, Languages.
Exposed through a small FastAPI app and secured by an admin API key. The full
output shape is in src/schemas.py.
Setup #
Requirements: Python 3.14 and PDM.
git clone git@tangled.org:did:plc:26n46rcm5csiruar6lrzf2fe tross && cd tross
pdm install
cp config.example.toml config.toml # then edit it
Configuration #
All config lives in config.toml at the project root, loaded by
src/config.py. Environment variables override the file.
The secondary tross key is there becuase it's personal account's cookies and misuse might lead to a ban.
admin_key = "..." # primary key, never expires
tross_key = "..." # secondary key
tross_key_valid_until = 2026-09-01 # rejected from Sept 2 onward
[linkedin]
[linkedin.cookies]
li_at = "AQED..." # required
JSESSIONID = '"ajax:..."' # required (keep the quotes)
Click on any profile while your network tab is open. You can right click the
call made right after the click and then copy it as bash. Then head to
curlconvertor and convert it to python's request. The cookie dict should be
as toml dict. The app derives the csrf-token from JSESSIONID at startup.
The app refuses to start if admin_key or tross_key is empty, or if li_at
or JSESSIONID is missing.
Running #
Local:
uvicorn app:app
Docker (mounts config.toml read-only, secrets never enter the image):
docker compose up -d --build
Swagger is at /docs, ReDoc is at /redoc (show redoc some love)
health endpoint is at /health
API #
GET /api/v1/scrape/ #
Scrape a LinkedIn profile.
- Query:
url(ahttps://www.linkedin.com/in/<vanity>/URL) - Header:
X-Admin-Key(a validadmin_keyortross_key)
curl -H "X-Admin-Key: $ADMIN_KEY" \
"http://localhost:8000/api/v1/scrape/?url=https://www.linkedin.com/in/heli-v-48b726106/"
Returns 200 with a Profile JSON object.
Errors use a structured body {"error_code": "...", "message": "..."}:
| Status | error_code | When |
|---|---|---|
| 401 | UNAUTHORIZED | Missing, invalid, or expired key |
| 400 | BAD_REQUEST | Malformed request (service layer) |
| 404 | NOT_FOUND | Resource not found (service layer) |
| 503 | SERVICE_UNAVAILABLE | Transient upstream failure (service) |
| 500 | none | Unexpected error (LinkedIn 5xx, parse) |
GET /health #
Unauthenticated liveness probe. Returns {"status": "ok"}.
Approach #
LinkedIn's profile page is an RSC stream: chunkId:JSON lines with $Lxx /
$xx references between chunks. The browser loads the full sections as
separate async POSTs to /flagship-web/rsc-action/actions/component. The
scraper copies that flow:
GET /flagship-web/in/<vanity>/ -> top card (name, headline, images)
POST /rsc-action/.../component x5 -> about, experience, education+certs,
languages, skills (fetched concurrently)
parse RSC stream -> resolve refs -> structure items -> Profile
Steps: parse chunks, resolve references, walk the React tree for text and
images, group flat text into items by component structure, then structure each
section into typed objects and validate as a Profile.
Hard parts and fixes:
- TLS fingerprinting:
httpxis rejected on the POST endpoints (GET 200, POST 500). The scraper usescurl_cffiwithimpersonate="firefox"to present a real browser TLS fingerprint. - Reverse-engineered POSTs: needs
componentId/sduiid/parentSpanIdparams, per-component payloads,csrf-token = JSESSIONID(quotes stripped), and page-instance headers copied from the top card. - Duplicate chunks: the stream re-emits the same content in several root
chunks.
_dedupe_repeatedcollapses exact block repetitions. Sections that repeat field values across items (certifications share issuer and issued date, languages can share proficiency) are de-triplicated, not value-deduplicated. - Flat text to items: a
componentKeystack groups items (wrapper at depth 1, items at depth 2). Certifications are split on their terminal field line (Credential ID, orIssuedwhen there is no credential id).
Known limitations #
- Multi-role experience: LinkedIn nests roles under one company. The scraper
returns one entry with the company in
titleand roles indescription. - Self-view profiles: some sections may be empty or show "Add ..." placeholders.
- Certification ordering: the issuer is the line after the name; rare
reorderings can mislabel it. Certs without a
Credential IDsplit onIssuedinstead. - Single image per field:
profile_imageandbackground_imagereturn one rendition URL, not all size variants. - No batching or rate-limiting or backoff.
- Cookies: if
li_atexpires, requests fail. There is no auto re-login (LinkedIn blocks it and it risks the account). - No caching: every request re-fetches from LinkedIn.
- Best-effort parsing: if LinkedIn changes the RSC layout, the parser may need adjustment.