This repository has no description
tross-assignment README.md
6.4 kB
Markdown
at main

Tross Assignment - LinkedIn Scraper #

Update #

I wasn't selected for the role. Here's the full email I got --

Hi Nikhil,

Thank you for taking the time to apply to Tross and complete the assignment. We appreciate the effort you put into the process.

We’ve now concluded this round of engineering hiring. For this role, we were looking for a specific mix of reverse engineering and software engineering experience, and have decided to move forward with a candidate whose experience aligns with our current requirements.

As we continue to grow our engineering team, there may be opportunities in the future where your profile could be a strong fit. We’d be happy to keep your application in consideration for upcoming roles.

Thank you again for your interest in Tross, and we wish you all the best.

Best,
Tross Careers Team

Async service that scrapes a public LinkedIn profile and returns it as a validated Pydantic Profile. It reads LinkedIn's RSC (React Server Components) wire format, not HTML, and rebuilds the full profile: top card plus About, Experience, Education, Skills, Certifications, Languages.

Exposed through a small FastAPI app and secured by an admin API key. The full output shape is in src/schemas.py.

Setup #

Requirements: Python 3.14 and PDM.

git clone git@tangled.org:did:plc:26n46rcm5csiruar6lrzf2fe tross && cd tross
pdm install
cp config.example.toml config.toml   # then edit it

Configuration #

All config lives in config.toml at the project root, loaded by src/config.py. Environment variables override the file.

The secondary tross key is there becuase it's personal account's cookies and misuse might lead to a ban.

admin_key = "..."                       # primary key, never expires
tross_key = "..."                       # secondary key
tross_key_valid_until = 2026-09-01      # rejected from Sept 2 onward

[linkedin]
[linkedin.cookies]
li_at      = "AQED..."                  # required
JSESSIONID = '"ajax:..."'               # required (keep the quotes)

Click on any profile while your network tab is open. You can right click the call made right after the click and then copy it as bash. Then head to curlconvertor and convert it to python's request. The cookie dict should be as toml dict. The app derives the csrf-token from JSESSIONID at startup.

The app refuses to start if admin_key or tross_key is empty, or if li_at or JSESSIONID is missing.

Running #

Local:

uvicorn app:app

Docker (mounts config.toml read-only, secrets never enter the image):

docker compose up -d --build

Swagger is at /docs, ReDoc is at /redoc (show redoc some love) health endpoint is at /health

API #

GET /api/v1/scrape/ #

Scrape a LinkedIn profile.

  • Query: url (a https://www.linkedin.com/in/<vanity>/ URL)
  • Header: X-Admin-Key (a valid admin_key or tross_key)
curl -H "X-Admin-Key: $ADMIN_KEY" \
  "http://localhost:8000/api/v1/scrape/?url=https://www.linkedin.com/in/heli-v-48b726106/"

Returns 200 with a Profile JSON object.

Errors use a structured body {"error_code": "...", "message": "..."}:

Status error_code When
401 UNAUTHORIZED Missing, invalid, or expired key
400 BAD_REQUEST Malformed request (service layer)
404 NOT_FOUND Resource not found (service layer)
503 SERVICE_UNAVAILABLE Transient upstream failure (service)
500 none Unexpected error (LinkedIn 5xx, parse)

GET /health #

Unauthenticated liveness probe. Returns {"status": "ok"}.

Approach #

LinkedIn's profile page is an RSC stream: chunkId:JSON lines with $Lxx / $xx references between chunks. The browser loads the full sections as separate async POSTs to /flagship-web/rsc-action/actions/component. The scraper copies that flow:

GET  /flagship-web/in/<vanity>/      -> top card (name, headline, images)
POST /rsc-action/.../component  x5  -> about, experience, education+certs,
                                       languages, skills (fetched concurrently)
parse RSC stream  ->  resolve refs  ->  structure items  ->  Profile

Steps: parse chunks, resolve references, walk the React tree for text and images, group flat text into items by component structure, then structure each section into typed objects and validate as a Profile.

Hard parts and fixes:

  • TLS fingerprinting: httpx is rejected on the POST endpoints (GET 200, POST 500). The scraper uses curl_cffi with impersonate="firefox" to present a real browser TLS fingerprint.
  • Reverse-engineered POSTs: needs componentId / sduiid / parentSpanId params, per-component payloads, csrf-token = JSESSIONID (quotes stripped), and page-instance headers copied from the top card.
  • Duplicate chunks: the stream re-emits the same content in several root chunks. _dedupe_repeated collapses exact block repetitions. Sections that repeat field values across items (certifications share issuer and issued date, languages can share proficiency) are de-triplicated, not value-deduplicated.
  • Flat text to items: a componentKey stack groups items (wrapper at depth 1, items at depth 2). Certifications are split on their terminal field line (Credential ID, or Issued when there is no credential id).

Known limitations #

  • Multi-role experience: LinkedIn nests roles under one company. The scraper returns one entry with the company in title and roles in description.
  • Self-view profiles: some sections may be empty or show "Add ..." placeholders.
  • Certification ordering: the issuer is the line after the name; rare reorderings can mislabel it. Certs without a Credential ID split on Issued instead.
  • Single image per field: profile_image and background_image return one rendition URL, not all size variants.
  • No batching or rate-limiting or backoff.
  • Cookies: if li_at expires, requests fail. There is no auto re-login (LinkedIn blocks it and it risks the account).
  • No caching: every request re-fetches from LinkedIn.
  • Best-effort parsing: if LinkedIn changes the RSC layout, the parser may need adjustment.