diff --git a/README.md b/README.md index 7993968..5816a7a 100644 --- a/README.md +++ b/README.md @@ -37,91 +37,9 @@ go run . search -did did:plc:abc123 -rkey 3abc # find si go run . search -did did:plc:abc123 -rkey 3abc -cluster # find similar + cluster ``` -### Search API - -`POST /api/v1/search` — bearer token optional. - -**Request:** -```json -{ - "query": "machine learning", - "did": "did:plc:abc123", - "rkey": "3abc", - "limit": 100, - "distinct": true, - "cluster": true, - "include_embeddings": true -} -``` - -| Field | Type | Default | Description | -|----------------------|--------|---------|-------------| -| `query` | string | — | Search query text (exclusive with `rkey`) | -| `did` | string | — | Scope search to a specific account (AT Protocol DID) | -| `rkey` | string | — | Use this post's embedding as query (requires `did`, exclusive with `query`) | -| `limit` | int | 400/1200 | Max results. Global: max 400. DID-scoped: max 1200 | -| `distinct` | bool | true | One result per account | -| `cluster` | bool | false | Enable UMAP+HDBSCAN clustering with c-TF-IDF topic extraction | -| `include_embeddings` | bool | false | Include per-result cluster embeddings (128d non-bearer, 768d bearer) | - -**Search modes:** At least one of `query`, `did`, or `rkey` is required. - -| Mode | Fields | Behavior | -|------|--------|----------| -| Global search | `query` | Semantic search across all posts (max 400) | -| Browse account | `did` | Newest posts, always clustered (max 1200) | -| Search in account | `did` + `query` | Semantic search within account's posts (max 1200) | -| Similar posts | `did` + `rkey` | Use post's embedding as query for global search | - -**Response:** -```json -{ - "results": [ - { - "did": "did:plc:abc123", - "handle": "user.bsky.social", - "collection": "app.bsky.feed.post", - "rkey": "3abc", - "text": "Post text...", - "score": -0.82, - "created_at": "2025-01-15T10:30:00Z", - "detected_lang": "en", - "cluster_id": 0, - "topics": ["topic1", "topic2"], - "embedding": [0.12, -0.34, ...] - } - ], - "clusters": [ - { - "id": 0, - "size": 15, - "topics": ["topic1", "topic2", "topic3"], - "result_indices": [0, 3, 7, 12] - } - ] -} -``` - -`clusters` and per-result `cluster_id`/`topics` are only present when `cluster: true`. `embedding` only present when `include_embeddings: true`. Score is negative inner product (lower = more similar). - -## Firehose protocol - -Zstd-compressed NDJSON over HTTP. Each line is a columnar batch: - -```json -{"did":["did:plc:abc","did:plc:xyz"],"col":["app.bsky.feed.post","app.bsky.actor.profile"],"rkey":["3abc","self"],"lang":["en","de"],"c":[[128 floats],[128 floats]],"r":[[128 floats],[128 floats]]} -``` - -| Field | Description | -|--------|-------------| -| `did` | AT Protocol DID (e.g. `did:plc:abc123`) | -| `col` | AT Protocol collection NSID (e.g. `app.bsky.feed.post`, `app.bsky.actor.profile`) | -| `rkey` | Record key within the collection | -| `lang` | Detected language (`en`, `de`) | -| `c` | Cluster embedding (prefix `"task: clustering \| query: "`) | -| `r` | Retrieval embedding (prefix `"title: none \| text: "`) | +### API documentation -Empty batches (`[]`) are heartbeats sent every 10s. +Full API spec (search endpoints, firehose protocol, schemas): **[OpenAPI 3.1](https://divepool.social/api/v1/openapi.json)** ## Verify embeddings diff --git a/experiments/requirements.txt b/experiments/requirements.txt new file mode 100644 index 0000000..f107310 --- /dev/null +++ b/experiments/requirements.txt @@ -0,0 +1,5 @@ +numpy>=1.26.0 +zstandard>=0.20.0 +scikit-learn>=1.3.0 +umap-learn>=0.5.0 +hdbscan>=0.8.33