* architecture brainstorming - 6/22 :ai:claude: ** multi-pattern episode title tagging :ai:claude: Problem: podcast episode titles are inconsistent and chaotic, making single regex patterns too brittle. Examples from Critical Role: - "Critical Role Campaign 2, Episode 2 - The Midnight Chase" - "Critical Role C2E3 - The Midnight Chase" - "CR C2E4: The Midnight Chase" - "Between the Sheets with Ashley Johnson" - "Talks Machina #47: Discussing Campaign 2, Episode 2" - "Critical Role Campaign 2 Episode 5 (FIXED AUDIO)" *** proposed approach: multi-pattern + llm improvement **** feed-specific pattern collections #+BEGIN_SRC typescript interface FeedPattern { feedId: string patterns: Array<{ name: string regex: string tagTemplate: string[] priority: number }> } // Example for Critical Role feed { patterns: [ { name: "Long Form Campaign Episodes", regex: "Campaign (\\d+).*Episode (\\d+)", tagTemplate: ["campaign:{1}", "episode:{2}", "type:main"], priority: 1 }, { name: "Abbreviated Format", regex: "C(\\d+)E(\\d+)", tagTemplate: ["campaign:{1}", "episode:{2}", "type:main"], priority: 2 }, { name: "Talks Machina", regex: "Talks Machina", tagTemplate: ["type:bonus", "format:discussion"], priority: 3 }, { name: "Interview Series", regex: "Between the Sheets", tagTemplate: ["type:bonus", "format:interview"], priority: 4 } ] } #+END_SRC **** llm-assisted pattern discovery Instead of training NER models, use LLM to discover new patterns when existing ones fail: #+BEGIN_SRC typescript interface TagExtractionPipeline { patterns: FeedPattern[] async improvePatterns(failedTitles: string[]): Promise { const prompt = ` I need regex patterns to extract tags from these podcast episode titles. Generate multiple patterns that handle the variations: Failed to match: ${failedTitles.join('\n- ')} Existing patterns: ${this.patterns.map(p => `- ${p.regex} → ${p.tagTemplate}`).join('\n')} Suggest additional patterns as JSON with regex, tagTemplate, and descriptive name. ` return await llm.generatePatterns(prompt) } async tagDirectly(title: string, feedContext: FeedContext): Promise { // Fallback: LLM directly tags weird edge cases } } #+END_SRC **** benefits of this approach - **Regex speed**: Fast pattern matching for 90% of episodes - **LLM intelligence**: Handles edge cases and discovers new patterns - **Incremental improvement**: System learns from failures - **Cost control**: Generate patterns once, not per episode - **User oversight**: Patterns can be reviewed/edited before adoption **** batch processing workflow 1. New feed subscribed → LLM analyzes first 20-50 episodes → suggests initial patterns 2. User reviews/approves/edits patterns 3. Ongoing episodes processed with regex patterns (fast) 4. Failed matches collected → periodic LLM pattern improvement 5. New patterns suggested → user review cycle This gives you both automation and consistency without the brittleness of single-pattern approaches or the complexity of training domain-specific NER models. ** dexie sync solutions analysis :ai:claude: *** existing dexie sync options **** dexie.syncable (legacy but functional) - **Purpose**: Two-way sync with remote servers via custom protocols - **Key requirement**: Must use UUID primary keys (prefixed with $$) - **Protocol support**: WebSocket, Ajax, or custom ISyncProtocol implementations - **Architecture**: Client-server model, not designed for P2P mesh **** dexie cloud (recommended saas) - **Purpose**: Batteries-included sync solution with authentication - **Benefits**: Automatic conflict resolution, user management, eager sync - **Limitations**: SaaS-only, not suitable for P2P realm architecture - **Cost**: Managed service, may not align with self-hosted goals *** integrating dexie.syncable with realm mesh **** custom isyncprotocol for webrtc p2p Can implement a custom sync protocol that uses WebRTC data channels as transport: #+BEGIN_SRC typescript Dexie.Syncable.registerSyncProtocol("realm-p2p", { sync: function(context, url, options, baseRevision, syncedRevision, changes, partial, applyRemoteChanges, onChangesAccepted, onSuccess, onError) { // Use existing RealmConnection for WebRTC transport // Convert Dexie changes to realm broadcast messages // Handle conflict resolution using HLC timestamps } }); // Usage db.syncable.connect("realm-p2p", realmId, { realmConnection: myExistingRealmConnection }); #+END_SRC **** mapping dexie changes to realm events #+BEGIN_SRC typescript interface DexieChange { table: string key: any obj?: any // undefined for deletes type: 1|2|3 // CREATE=1, UPDATE=2, DELETE=3 } // Convert to realm-friendly events interface RealmSyncEvent { type: 'sync.changes' revision: number changes: DexieChange[] timestamp: HLC } #+END_SRC **** benefits of dexie.syncable integration - **Proven sync logic**: Dexie handles conflict detection, partial changes, revision tracking - **UUID keys**: Already designed for distributed systems - **Atomic operations**: Transactions ensure consistency during sync - **Change tracking**: Built-in hooks for detecting local modifications **** challenges with realm mesh integration - **Server-centric design**: ISyncProtocol assumes single remote endpoint, not mesh - **Revision model**: Dexie uses linear revisions, mesh needs vector clocks/HLC - **Conflict resolution**: Dexie's conflict model may not align with HLC-based resolution *** alternative approach: dexie observable + custom sync Instead of forcing Dexie.Syncable into P2P model, use Dexie's change observation: #+BEGIN_SRC typescript // Listen to local changes db.myTable.hook('creating', (primKey, obj, trans) => { // Convert to HLC event and broadcast to realm const event = createHLCEvent('create', 'myTable', obj) realmConnection.broadcast(event) }) // Apply remote changes realmConnection.addEventListener('peerdata', (event) => { if (event.data.type === 'sync.change') { // Apply change to local Dexie with conflict resolution await applyRemoteChange(event.data) } }) #+END_SRC **** recommendation For your realm-based P2P architecture, **custom sync with Dexie hooks** is probably cleaner than forcing Dexie.Syncable into a mesh model. You get: - Full control over HLC-based conflict resolution - Natural integration with existing realm WebRTC infrastructure - Simpler architecture without protocol translation layer - Ability to use CRDTs for specific data types However, you lose Dexie.Syncable's battle-tested sync logic and would need to implement revision tracking, partial sync, and conflict detection yourself. ** reactive database alternatives to dexie :ai:claude: *** comparison matrix: dexie vs rxdb vs tanstack db **** dexie (current choice) - **Reactive queries**: liveQuery() function, binary range tree for efficient change detection - **API simplicity**: Minimalistic IndexedDB wrapper, straightforward to learn - **Performance**: Fast for most use cases, some reports of slowness >100 entries in reactive mode - **Sync options**: Dexie.Syncable (legacy) or Dexie Cloud (SaaS) - **Size**: Lightweight, small bundle - **Preact integration**: Works with any framework, manual integration needed **** rxdb - **Reactive queries**: Built on RxJS observables, deep reactivity integration - **API complexity**: MongoDB-like queries, JSON Schema validation, more features - **Performance**: Claims superior performance, especially with premium IndexedDB storage - **Sync options**: WebRTC P2P built-in, CouchDB replication, custom adapters - **Size**: Larger bundle, enterprise-grade features - **Preact integration**: Can use Dexie as storage layer + RxDB features on top **** tanstack db (alpha) - **Reactive queries**: Sub-millisecond live queries via differential dataflow - **API approach**: Collections + fine-grained reactivity, not traditional ORM - **Performance**: "Blazing fast" claims, incremental query updates - **Sync options**: Backend-agnostic, designed for sync engines like ElectricSQL - **Size**: Unknown, alpha stage - **Preact integration**: Builds on TanStack Query + Preact Signals *** preact-native reactive solutions **** preact signals + manual persistence #+BEGIN_SRC typescript import { signal, computed } from '@preact/signals' // Reactive podcast state const podcasts = signal([]) const playState = signal({}) // Computed derived state const unplayedEpisodes = computed(() => podcasts.value.filter(p => !playState.value[p.id]?.completed) ) // Manual persistence layer const persistToDB = async (table, data) => { // Could be Dexie, IndexedDB, or any storage await db[table].put(data) } // Reactive sync const syncChanges = (change) => { realmConnection.broadcast(change) } #+END_SRC **Benefits**: - Native Preact integration - Full control over sync logic - Lightweight, exactly what you need **Downsides**: - Manual implementation of persistence, querying, indexing **** zustand + tanstack query + preact signals #+BEGIN_SRC typescript import { create } from 'zustand' import { useQuery } from '@tanstack/react-query' import { signal } from '@preact/signals' // Local state management const usePodcastStore = create((set) => ({ subscriptions: [], playState: {}, addSubscription: (feed) => set(state => ({ subscriptions: [...state.subscriptions, feed] })) })) // Server state + sync const useEpisodes = (feedId) => useQuery({ queryKey: ['episodes', feedId], queryFn: () => fetchEpisodesFromPeers(feedId) }) // Reactive UI state const currentEpisode = signal(null) #+END_SRC **Benefits**: - Proven patterns, excellent ecosystem - Separation of local vs server state - Great Preact integration via @preact-signals/query **Downsides**: - More complex setup - Multiple state libraries to coordinate *** recommendation for skypod architecture Given your P2P realm-based sync and Preact frontend: **Option 1: Preact Signals + Simple Persistence** (Recommended for MVP) - Use Preact Signals for reactive state management - Simple IndexedDB wrapper (maybe tiny-indexeddb) for persistence - Custom HLC-based sync via your existing realm infrastructure - Minimal dependencies, full control **Option 2: RxDB with WebRTC** (If you want battle-tested sync) - RxDB handles the complex sync logic - Built-in WebRTC P2P support aligns with your architecture - More dependencies but proven in production **Option 3: Wait for TanStack DB** (Future consideration) - Promising architecture for reactive local-first apps - Still alpha, but could be perfect fit when stable - Would integrate beautifully with your existing TanStack ecosystem **** why not stick with dexie? - Dexie is solid, but its sync model doesn't align well with P2P mesh - Your existing realm WebRTC infrastructure is already doing the hard work - Simpler to use Dexie just for storage + build reactive layer on top For your use case, **Preact Signals + manual persistence** might be the sweet spot - gives you exactly what you need without fighting impedance mismatches. ** realm-integrated storage manager design :ai:claude: *** storage manager with table definitions Design a storage manager that integrates directly with RealmConnection and handles table schemas declaratively: #+BEGIN_SRC typescript interface TableSchema { name: string schema: string // Dexie schema string like '++id, feedUrl, title' indexes?: string[] } class RealmStorageManager { private db: Dexie private signals = new Map>() private schemas = new Map() constructor(private realmConnection: RealmConnection) { this.db = new Dexie('SkypodDB') this.setupRealmSync() } // Create reactive signal with automatic table setup createSignal(tableSchema: TableSchema, initialValue: T[] = []): Signal { const { name, schema } = tableSchema // Store schema for DB setup this.schemas.set(name, tableSchema) // Create reactive signal const signal = signal(initialValue) this.signals.set(name, signal) // Setup table in Dexie (deferred until all schemas defined) this.ensureTableExists(name, schema) // Load initial data from storage this.loadInitialData(name, signal) return signal } private ensureTableExists(tableName: string, schema: string) { // Rebuild DB schema with all known tables const allSchemas = Object.fromEntries( Array.from(this.schemas.entries()).map(([name, {schema}]) => [name, schema]) ) this.db.close() this.db = new Dexie('SkypodDB') this.db.version(1).stores(allSchemas) } private async loadInitialData(tableName: string, signal: Signal) { try { const data = await this.db[tableName].toArray() signal.value = data } catch (err) { console.warn(`Failed to load ${tableName}:`, err) } } // CRUD operations that auto-sync and persist async add(tableName: string, item: Omit): Promise { const id = nanoid() const fullItem = { ...item, id } as T // Persist to storage await this.db[tableName].put(fullItem) // Update signal const signal = this.signals.get(tableName) signal.value = [...signal.value, fullItem] // Broadcast to realm this.realmConnection.broadcast({ type: 'storage.change', table: tableName, operation: 'add', data: fullItem, hlc: this.generateHLC() }) return id } async update(tableName: string, id: string, changes: Partial): Promise { // Persist to storage await this.db[tableName].update(id, changes) // Update signal const signal = this.signals.get(tableName) signal.value = signal.value.map(item => item.id === id ? { ...item, ...changes } : item ) // Broadcast to realm this.realmConnection.broadcast({ type: 'storage.change', table: tableName, operation: 'update', data: { id, changes }, hlc: this.generateHLC() }) } async remove(tableName: string, id: string): Promise { // Remove from storage await this.db[tableName].delete(id) // Update signal const signal = this.signals.get(tableName) signal.value = signal.value.filter(item => item.id !== id) // Broadcast to realm this.realmConnection.broadcast({ type: 'storage.change', table: tableName, operation: 'remove', data: { id }, hlc: this.generateHLC() }) } // Handle incoming changes from realm peers private setupRealmSync() { this.realmConnection.addEventListener('peerdata', (event) => { if (event.data.type === 'storage.change') { this.applyRemoteChange(event.data) } }) } private async applyRemoteChange(change: any) { const { table, operation, data } = change const signal = this.signals.get(table) if (!signal) return // Table not initialized on this peer switch (operation) { case 'add': // Check for conflicts/duplicates with HLC resolution if (!signal.value.find(item => item.id === data.id)) { signal.value = [...signal.value, data] await this.db[table].put(data) } break case 'update': signal.value = signal.value.map(item => item.id === data.id ? { ...item, ...data.changes } : item ) await this.db[table].update(data.id, data.changes) break case 'remove': signal.value = signal.value.filter(item => item.id !== data.id) await this.db[table].delete(data.id) break } } private generateHLC(): string { // Use your HLC implementation return `${Date.now()}.0.${this.realmConnection.identid}` } } #+END_SRC *** usage in components #+BEGIN_SRC typescript // Setup storage manager const storage = new RealmStorageManager(realmConnection) // Define reactive data stores with schemas const podcasts = storage.createSignal({ name: 'podcasts', schema: '++id, feedUrl, title, description, imageUrl, tags' }) const episodes = storage.createSignal({ name: 'episodes', schema: '++id, podcastId, guid, title, audioUrl, publishedAt, duration' }) const playState = storage.createSignal({ name: 'playState', schema: '++id, episodeId, deviceId, position, completed, updatedAt' }) // Use in components - automatically reactive function PodcastList() { return (
{podcasts.value.map(podcast => ( ))}
) } function PodcastItem({ podcast }) { const handleSubscribe = async () => { await storage.add('podcasts', { feedUrl: 'https://example.com/feed.xml', title: 'New Podcast', tags: ['tech', 'programming'] }) } const handleUpdateTags = async () => { await storage.update('podcasts', podcast.id, { tags: [...podcast.tags, 'favorite'] }) } return (

{podcast.title}

) } #+END_SRC *** benefits of this approach - **Declarative schemas**: Table definitions alongside signal creation - **Automatic persistence**: All changes auto-saved to IndexedDB - **Realm sync integration**: Changes broadcast and applied via existing WebRTC mesh - **Type safety**: Generic createSignal() for TypeScript support - **Reactive UI**: Preact components automatically re-render on data changes - **Conflict resolution**: HLC timestamps for distributed conflict handling - **Simple API**: Just add/update/remove, sync handled automatically This gives you a clean abstraction where components just work with reactive signals, while persistence and P2P sync happen transparently in the background. ** idb + signals storage architecture :ai:claude: *** data model for podcast app **** core entities - **Feed**: RSS/podcast source with metadata - **Entry**: Individual episodes with content + auto-generated tags - **PlayState**: Per-device playback position/completion status - **Subscription**: User's relationship to feeds (folder path, settings) - **Filter**: Saved query/view configurations **** schema design #+BEGIN_SRC typescript interface Feed { id: string // nanoid url: string // RSS feed URL (unique) title: string description?: string imageUrl?: string language?: string lastFetched?: number etag?: string // HTTP caching refreshInterval?: number } interface Entry { id: string // nanoid feedId: string // reference to Feed guid: string // RSS guid (unique per feed) title: string description?: string audioUrl?: string imageUrl?: string publishedAt: number // timestamp duration?: number // seconds fileSize?: number // bytes // Auto-generated tags from title parsing tags: string[] // ["series:criticalrole", "episode:84", "type:main"] } interface PlayState { id: string // nanoid entryId: string // reference to Entry deviceId: string // this device's identity position: number // playback position in seconds completed: boolean // fully played speed: number // playback speed updatedAt: number // HLC timestamp for sync } interface Subscription { id: string // nanoid feedId: string // reference to Feed folderPath: string // "/Tech/Programming" or "" for root autoDownload: boolean notificationsEnabled: boolean createdAt: number } interface Filter { id: string // nanoid name: string // "Critical Role Campaign 2" folderPath: string // where to show in sidebar "/Channels/Critical Role" query: FilterQuery // what to include/exclude isDefault: boolean // auto-created for new feeds position: number // ordering in UI } interface FilterQuery { includeTags?: string[] // must have ALL these tags excludeTags?: string[] // must NOT have any of these feedIds?: string[] // limit to specific feeds completed?: boolean // filter by play state dateRange?: [number, number] // publishedAt range } #+END_SRC *** idb + signals storage manager #+BEGIN_SRC typescript import { openDB } from 'idb' import { signal, computed } from '@preact/signals' import { nanoid } from 'nanoid' class PodcastStorage { // Core data signals feeds = signal([]) entries = signal([]) playStates = signal([]) subscriptions = signal([]) filters = signal([]) // Loading states loaded = signal(false) private db: IDBDatabase private deviceId: string constructor(private realmConnection: RealmConnection, deviceId: string) { this.deviceId = deviceId this.setupIDB() this.setupSync() } // Computed queries (reactive!) subscribedFeeds = computed(() => { const subFeedIds = new Set(this.subscriptions.value.map(s => s.feedId)) return this.feeds.value.filter(f => subFeedIds.has(f.id)) }) entriesForFeed = computed(() => (feedId: string) => this.entries.value.filter(e => e.feedId === feedId) ) unplayedEntries = computed(() => { const playStateMap = new Map( this.playStates.value .filter(ps => ps.deviceId === this.deviceId) .map(ps => [ps.entryId, ps]) ) return this.entries.value.filter(e => !playStateMap.get(e.id)?.completed ) }) entriesForFilter = computed(() => (filter: Filter) => { let entries = this.entries.value // Filter by feeds if (filter.query.feedIds?.length) { const feedIds = new Set(filter.query.feedIds) entries = entries.filter(e => feedIds.has(e.feedId)) } // Filter by tags if (filter.query.includeTags?.length) { entries = entries.filter(e => filter.query.includeTags.every(tag => e.tags.includes(tag)) ) } if (filter.query.excludeTags?.length) { entries = entries.filter(e => !filter.query.excludeTags.some(tag => e.tags.includes(tag)) ) } // Filter by completion status if (filter.query.completed !== undefined) { const playStateMap = new Map( this.playStates.value .filter(ps => ps.deviceId === this.deviceId) .map(ps => [ps.entryId, ps]) ) entries = entries.filter(e => { const isCompleted = playStateMap.get(e.id)?.completed ?? false return isCompleted === filter.query.completed }) } return entries.sort((a, b) => b.publishedAt - a.publishedAt) }) // Database setup private async setupIDB() { this.db = await openDB('SkypodDB', 1, { upgrade(db) { // Feeds store const feedStore = db.createObjectStore('feeds', { keyPath: 'id' }) feedStore.createIndex('url', 'url', { unique: true }) // Entries store const entryStore = db.createObjectStore('entries', { keyPath: 'id' }) entryStore.createIndex('feedId', 'feedId') entryStore.createIndex('publishedAt', 'publishedAt') entryStore.createIndex('tags', 'tags', { multiEntry: true }) // PlayStates store const playStateStore = db.createObjectStore('playStates', { keyPath: 'id' }) playStateStore.createIndex('entryId', 'entryId') playStateStore.createIndex('deviceId', 'deviceId') playStateStore.createIndex('entryDevice', ['entryId', 'deviceId'], { unique: true }) // Subscriptions store const subStore = db.createObjectStore('subscriptions', { keyPath: 'id' }) subStore.createIndex('feedId', 'feedId', { unique: true }) // Filters store const filterStore = db.createObjectStore('filters', { keyPath: 'id' }) filterStore.createIndex('position', 'position') } }) await this.loadAllData() } private async loadAllData() { const tx = this.db.transaction(['feeds', 'entries', 'playStates', 'subscriptions', 'filters'], 'readonly') const [feeds, entries, playStates, subscriptions, filters] = await Promise.all([ tx.objectStore('feeds').getAll(), tx.objectStore('entries').getAll(), tx.objectStore('playStates').getAll(), tx.objectStore('subscriptions').getAll(), tx.objectStore('filters').getAll() ]) this.feeds.value = feeds this.entries.value = entries this.playStates.value = playStates this.subscriptions.value = subscriptions this.filters.value = filters this.loaded.value = true } // CRUD operations with sync async addFeed(feedData: Omit): Promise { const feed: Feed = { ...feedData, id: nanoid() } // Persist const tx = this.db.transaction('feeds', 'readwrite') await tx.objectStore('feeds').add(feed) // Update signal this.feeds.value = [...this.feeds.value, feed] // Sync to realm this.realmConnection.broadcast({ type: 'storage.change', table: 'feeds', operation: 'add', data: feed, hlc: this.generateHLC() }) return feed.id } async updatePlayState(entryId: string, updates: Partial>): Promise { // Find existing or create new let playState = this.playStates.value.find(ps => ps.entryId === entryId && ps.deviceId === this.deviceId ) if (playState) { playState = { ...playState, ...updates, updatedAt: Date.now() } // Update in IDB const tx = this.db.transaction('playStates', 'readwrite') await tx.objectStore('playStates').put(playState) // Update signal this.playStates.value = this.playStates.value.map(ps => ps.id === playState.id ? playState : ps ) } else { playState = { id: nanoid(), entryId, deviceId: this.deviceId, position: 0, completed: false, speed: 1.0, updatedAt: Date.now(), ...updates } // Add to IDB const tx = this.db.transaction('playStates', 'readwrite') await tx.objectStore('playStates').add(playState) // Update signal this.playStates.value = [...this.playStates.value, playState] } // Sync to realm this.realmConnection.broadcast({ type: 'storage.change', table: 'playStates', operation: playState ? 'update' : 'add', data: playState, hlc: this.generateHLC() }) } async createFilter(filterData: Omit): Promise { const filter: Filter = { ...filterData, id: nanoid() } // Persist const tx = this.db.transaction('filters', 'readwrite') await tx.objectStore('filters').add(filter) // Update signal this.filters.value = [...this.filters.value, filter] // Sync to realm this.realmConnection.broadcast({ type: 'storage.change', table: 'filters', operation: 'add', data: filter, hlc: this.generateHLC() }) return filter.id } // Sync handling private setupSync() { this.realmConnection.addEventListener('peerdata', (event) => { if (event.data.type === 'storage.change') { this.applyRemoteChange(event.data) } }) } private async applyRemoteChange(change: any) { const { table, operation, data } = change switch (table) { case 'feeds': await this.applyFeedChange(operation, data) break case 'entries': await this.applyEntryChange(operation, data) break case 'playStates': await this.applyPlayStateChange(operation, data) break case 'subscriptions': await this.applySubscriptionChange(operation, data) break case 'filters': await this.applyFilterChange(operation, data) break } } private async applyPlayStateChange(operation: string, data: PlayState) { switch (operation) { case 'add': // Check for duplicates if (!this.playStates.value.find(ps => ps.id === data.id)) { await this.db.transaction('playStates', 'readwrite').objectStore('playStates').add(data) this.playStates.value = [...this.playStates.value, data] } break case 'update': // Conflict resolution: latest updatedAt wins const existing = this.playStates.value.find(ps => ps.entryId === data.entryId && ps.deviceId === data.deviceId ) if (!existing || data.updatedAt > existing.updatedAt) { await this.db.transaction('playStates', 'readwrite').objectStore('playStates').put(data) this.playStates.value = this.playStates.value.map(ps => ps.entryId === data.entryId && ps.deviceId === data.deviceId ? data : ps ) } break } } private generateHLC(): string { // Implement your HLC logic return `${Date.now()}.0.${this.deviceId}` } } #+END_SRC *** usage in components #+BEGIN_SRC typescript // App setup const storage = new PodcastStorage(realmConnection, deviceId) // Component usage - automatically reactive function PodcastSidebar() { const filters = storage.filters.value return (
{filters.map(filter => ( ))}
) } function FilterView({ filter, storage }) { const entries = storage.entriesForFilter.value(filter) return (

{filter.name} ({entries.length})

{entries.map(entry => ( ))}
) } function EntryItem({ entry, storage }) { const playState = storage.playStates.value.find(ps => ps.entryId === entry.id && ps.deviceId === storage.deviceId ) const handlePlay = async () => { await storage.updatePlayState(entry.id, { position: 0, updatedAt: Date.now() }) } const handleMarkCompleted = async () => { await storage.updatePlayState(entry.id, { completed: true, position: entry.duration || 0, updatedAt: Date.now() }) } return (

{entry.title}

Tags: {entry.tags.join(', ')}
{playState &&
Position: {playState.position}s
}
) } #+END_SRC *** benefits of this approach - **Lightweight**: ~2KB idb vs ~30KB Dexie - **Reactive queries**: Computed signals automatically update UI - **Flexible filtering**: Tag-based system with complex queries - **Conflict resolution**: HLC timestamps for play state sync - **Device-specific data**: Play states scoped per device - **Efficient**: Only re-compute affected queries when data changes - **Type safe**: Full TypeScript support throughout ** architectural summary and performance considerations :ai:claude: *** core architecture principles What we've settled on: 1. **Each client maintains IndexedDB as their source of truth** - Local storage is authoritative for that device - No external database dependency - Offline-first by design 2. **Storage manager provides reactive signals that auto-persist** - Signals automatically store deltas to IndexedDB - Changes published to realm peers via WebRTC - UI automatically updates via signal reactivity 3. **HLC timestamps for event ordering, not replay** - Events include HLC for distributed ordering - Used to ensure reducers apply changes in consistent order - NOT for full event sourcing replay - just conflict resolution - Example: two devices update same play position, latest HLC wins *** sync model: state-based with events This is **not** true event sourcing: - No event replay required - No need to rebuild state from event log - Events are just sync messages with conflict resolution Instead: - Materialized state stored in IndexedDB (feeds, episodes, play states) - Changes broadcast as events with HLC timestamps - Receivers apply events to their materialized state - Conflicts resolved via HLC comparison (latest wins) #+BEGIN_SRC typescript // Example: Play state conflict resolution private async applyPlayStateChange(operation: string, data: PlayState) { const existing = this.playStates.value.find(ps => ps.entryId === data.entryId && ps.deviceId === data.deviceId ) // Conflict resolution: latest HLC timestamp wins if (!existing || data.updatedAt > existing.updatedAt) { // Apply change to IndexedDB + signal await this.db.transaction('playStates', 'readwrite').objectStore('playStates').put(data) this.playStates.value = this.playStates.value.map(ps => ps.entryId === data.entryId && ps.deviceId === data.deviceId ? data : ps ) } // Else: ignore older change } #+END_SRC *** performance scaling strategy **** start simple (recommended for mvp) - All data in signals, computed queries via array operations - 100k play states = ~27 years of heavy podcast listening - Most apps won't hit performance issues for years **** optimization path when needed 1. **Pre-computed indexes** (5 minute fix) #+BEGIN_SRC typescript playStatesByEntry = computed(() => { const map = new Map() for (const ps of this.playStates.value) { if (ps.deviceId === this.deviceId) { map.set(ps.entryId, ps) } } return map }) #+END_SRC 2. **Lazy loading for old data** (keep working set small) #+BEGIN_SRC typescript // Only recent play states in signals recentPlayStates = signal([]) // Last 30 days // Lazy load from IDB for historical lookups async getPlayState(entryId: string): Promise { const recent = this.recentPlayStates.value.find(ps => ps.entryId === entryId) if (recent) return recent // Fall back to IDB query for older data return await this.db.transaction('playStates', 'readonly') .objectStore('playStates').index('entryDevice').get([entryId, this.deviceId]) } #+END_SRC 3. **UI pagination/virtualization** (only compute visible items) **** hybrid approach (if conservative about performance) - Keep heavy queries in IndexedDB using indexes - Use signals for lightweight UI state only - Manually trigger re-fetch when data changes *** why this architecture works - **Lightweight**: 2KB idb vs 30KB Dexie, no complex sync protocol - **Reactive**: Preact Signals provide efficient UI updates - **Conflict-free**: Device-specific data (play states) + HLC for shared data - **Scalable**: Can optimize computed queries without architectural changes - **Simple**: No event replay, no complex CRDT logic, just state + deltas - **Resilient**: Each device is self-contained, realm just for sync This gives you a local-first architecture that's reactive, performant, and scales naturally with your existing realm infrastructure. ** redux-style action dispatch architecture :ai:claude: *** redux + signals approach Better than method-based storage: explicit action dispatch to storage engine #+BEGIN_SRC typescript // Action types interface FeedActions { type: 'feeds/add' payload: Omit meta?: { skipSync?: boolean } } interface PlayStateActions { type: 'playStates/update' payload: { entryId: string; updates: Partial } meta?: { skipSync?: boolean } } interface FilterActions { type: 'filters/create' payload: Omit meta?: { skipSync?: boolean } } type StorageAction = FeedActions | PlayStateActions | FilterActions // Storage engine with reducer pattern class PodcastStorage { // Reactive state feeds = signal([]) entries = signal([]) playStates = signal([]) filters = signal([]) constructor(private realmConnection: RealmConnection, private deviceId: string) { this.setupIDB() this.setupRealmSync() } // Single dispatch method async dispatch(action: StorageAction): Promise { // 1. Apply to local state + IDB await this.reduce(action) // 2. Sync to realm (unless explicitly skipped) if (!action.meta?.skipSync) { this.realmConnection.broadcast({ type: 'storage.action', action, hlc: this.generateHLC() }) } } // Reducer handles all state changes private async reduce(action: StorageAction): Promise { switch (action.type) { case 'feeds/add': { const feed: Feed = { ...action.payload, id: nanoid() } // Persist to IDB await this.db.transaction('feeds', 'readwrite').objectStore('feeds').add(feed) // Update signal this.feeds.value = [...this.feeds.value, feed] break } case 'playStates/update': { const { entryId, updates } = action.payload let playState = this.playStates.value.find(ps => ps.entryId === entryId && ps.deviceId === this.deviceId ) if (playState) { playState = { ...playState, ...updates, updatedAt: Date.now() } await this.db.transaction('playStates', 'readwrite').objectStore('playStates').put(playState) this.playStates.value = this.playStates.value.map(ps => ps.id === playState.id ? playState : ps ) } else { playState = { id: nanoid(), entryId, deviceId: this.deviceId, position: 0, completed: false, speed: 1.0, updatedAt: Date.now(), ...updates } await this.db.transaction('playStates', 'readwrite').objectStore('playStates').add(playState) this.playStates.value = [...this.playStates.value, playState] } break } case 'filters/create': { const filter: Filter = { ...action.payload, id: nanoid() } await this.db.transaction('filters', 'readwrite').objectStore('filters').add(filter) this.filters.value = [...this.filters.value, filter] break } } } // Handle remote actions from realm peers private setupRealmSync() { this.realmConnection.addEventListener('peerdata', async (event) => { if (event.data.type === 'storage.action') { // Apply remote action without re-syncing const actionWithSkipSync = { ...event.data.action, meta: { ...event.data.action.meta, skipSync: true } } await this.dispatch(actionWithSkipSync) } }) } } #+END_SRC *** usage in components #+BEGIN_SRC typescript // Components dispatch actions instead of calling methods function PodcastItem({ podcast, storage }) { const handleSubscribe = async () => { await storage.dispatch({ type: 'feeds/add', payload: { url: 'https://example.com/feed.xml', title: 'New Podcast', description: 'A great show' } }) } const handleUpdatePlayState = async () => { await storage.dispatch({ type: 'playStates/update', payload: { entryId: currentEpisode.id, updates: { position: 120, completed: false } } }) } return (

{podcast.title}

) } #+END_SRC *** benefits of redux-style dispatch **** predictable state management - All state changes go through single `dispatch` method - Easy to debug: log all actions - Testable: pure reducer functions - Time-travel debugging possible **** perfect sync integration #+BEGIN_SRC typescript // Local action await storage.dispatch({ type: 'playStates/update', payload: { entryId: 'ep123', updates: { position: 150 } } }) // Remote action (automatically applied without re-sync) // Realm peer receives action and dispatches with skipSync: true #+END_SRC **** conflict resolution built-in #+BEGIN_SRC typescript // HLC conflict resolution in reducer case 'playStates/update': { const existing = this.playStates.value.find(ps => ps.entryId === action.payload.entryId && ps.deviceId === action.payload.deviceId ) // Only apply if newer (HLC comparison) if (!existing || action.payload.updatedAt > existing.updatedAt) { // Apply update } // Else: ignore older update break } #+END_SRC **** middleware support #+BEGIN_SRC typescript class PodcastStorage { private middleware: Middleware[] = [] async dispatch(action: StorageAction): Promise { // Run middleware pipeline let processedAction = action for (const mw of this.middleware) { processedAction = await mw(processedAction, this) } await this.reduce(processedAction) if (!processedAction.meta?.skipSync) { this.realmConnection.broadcast({ type: 'storage.action', action: processedAction, hlc: this.generateHLC() }) } } } // Example: logging middleware const loggingMiddleware: Middleware = (action, store) => { console.log('Dispatching action:', action.type, action.payload) return action } #+END_SRC *** comparison: methods vs redux dispatch **** method-based approach #+BEGIN_SRC typescript // Multiple methods to remember await storage.addFeed(feedData) await storage.updatePlayState(entryId, updates) await storage.createFilter(filterData) await storage.removeFeed(feedId) // Each method handles persistence + sync differently #+END_SRC **** redux dispatch approach #+BEGIN_SRC typescript // Single interface for all changes await storage.dispatch({ type: 'feeds/add', payload: feedData }) await storage.dispatch({ type: 'playStates/update', payload: { entryId, updates } }) await storage.dispatch({ type: 'filters/create', payload: filterData }) await storage.dispatch({ type: 'feeds/remove', payload: { id: feedId } }) // Consistent handling: all actions go through same pipeline #+END_SRC *** why redux-style is better for sync - **Consistent**: All state changes use same pattern - **Debuggable**: Easy to log and replay actions - **Testable**: Pure reducers, predictable state changes - **Sync-friendly**: Actions naturally map to sync messages - **Extensible**: Middleware for logging, validation, etc. - **Conflict resolution**: Built into reducer logic This gives you Redux's predictability with Preact Signals' reactivity and seamless P2P sync integration. ** reactive signal indices pattern :ai:claude: *** readonly signals + action dispatch architecture Storage engine vends readonly signals and only accepts changes via action dispatch: #+BEGIN_SRC typescript class PodcastStorage { // Read-only signals - components can't mutate directly readonly feeds: ReadonlySignal readonly entries: ReadonlySignal readonly playStates: ReadonlySignal readonly filters: ReadonlySignal // Private writable signals (internal only) private _feeds = signal([]) private _entries = signal([]) private _playStates = signal([]) private _filters = signal([]) constructor() { // Expose as readonly to prevent direct mutation this.feeds = this._feeds this.entries = this._entries this.playStates = this._playStates this.filters = this._filters } // Only way to change state is via dispatch async dispatch(action: StorageAction): Promise { await this.reduce(action) if (!action.meta?.skipSync) { this.broadcast(action) } } // Computed selectors - no manual slicing needed subscribedFeeds = computed(() => { const subFeedIds = new Set(this.subscriptions.value.map(s => s.feedId)) return this.feeds.value.filter(f => subFeedIds.has(f.id)) }) } #+END_SRC *** signal indices: reactive database-like queries Create cached computed signals that act like database indices: #+BEGIN_SRC typescript class PodcastStorage { // Source of truth - signal of arrays readonly entries: ReadonlySignal = this._entries readonly playStates: ReadonlySignal = this._playStates // Reactive indices - cached computed signals private episodeById = new Map>() private entriesForFeed = new Map>() private playStateForEntry = new Map>() // Get individual episode (like SELECT * FROM episodes WHERE id = ?) getEpisode(id: string): ReadonlySignal { if (!this.episodeById.has(id)) { const signal = computed(() => this.entries.value.find(e => e.id === id) || null ) this.episodeById.set(id, signal) } return this.episodeById.get(id)! } // Get episodes for feed (like SELECT * FROM episodes WHERE feedId = ?) getEntriesForFeed(feedId: string): ReadonlySignal { if (!this.entriesForFeed.has(feedId)) { const signal = computed(() => this.entries.value.filter(e => e.feedId === feedId) ) this.entriesForFeed.set(feedId, signal) } return this.entriesForFeed.get(feedId)! } // Get play state for episode (like JOIN query) getPlayStateForEntry(entryId: string): ReadonlySignal { if (!this.playStateForEntry.has(entryId)) { const signal = computed(() => this.playStates.value.find(ps => ps.entryId === entryId && ps.deviceId === this.deviceId ) || null ) this.playStateForEntry.set(entryId, signal) } return this.playStateForEntry.get(entryId)! } } #+END_SRC *** comparison to database indices | Database Index | Signal Index | |---|---| | `CREATE INDEX episodes_by_feed ON episodes(feedId)` | `getEntriesForFeed(feedId)` | | `CREATE INDEX playstate_by_entry ON playstates(entryId)` | `getPlayStateForEntry(entryId)` | | `SELECT * FROM episodes WHERE id = ?` | `getEpisode(id)` | *** automatic index maintenance #+BEGIN_SRC typescript // When you dispatch an action that changes episodes... await storage.dispatch({ type: 'entries/add', payload: newEpisode }) // All the relevant cached signals automatically update: // - getEpisode(newEpisode.id) now returns the new episode // - getEntriesForFeed(newEpisode.feedId) includes the new episode // - No manual cache invalidation needed! #+END_SRC *** component usage with granular reactivity #+BEGIN_SRC typescript function EpisodePage({ episodeId, storage }) { const episode = storage.getEpisode(episodeId).value const playState = storage.getPlayStateForEntry(episodeId).value // These only re-render when THIS specific episode/playstate changes // Not when any other episode in the entire array changes const handlePlay = () => { storage.dispatch({ type: 'playStates/update', payload: { entryId: episodeId, updates: { position: 0 } } }) } return (

{episode?.title}

Position: {playState?.position || 0}

) } function FeedPage({ feedId, storage }) { const episodes = storage.getEntriesForFeed(feedId).value // Only re-renders when episodes for THIS feed change // Not when episodes for other feeds change return (

Episodes ({episodes.length})

{episodes.map(episode => ( ))}
) } #+END_SRC *** benefits of signal indices pattern **** performance benefits - **Granular reactivity**: Only components using specific data re-render - **Efficient queries**: No scanning entire arrays repeatedly - **Automatic caching**: Computed signals cache results until dependencies change - **Lazy loading**: Indices created on-demand **** maintenance benefits - **Automatic invalidation**: Computed signals update when source data changes - **No manual cache management**: No need to invalidate caches manually - **Consistent state**: Indices always reflect current source of truth - **Single source of truth**: Array signals remain authoritative **** developer experience - **Database-like queries**: Familiar patterns for data access - **Type safety**: Full TypeScript support for query results - **Predictable**: All mutations go through action dispatch - **Debuggable**: Easy to track what data components are using *** final architecture summary 1. **Storage engine owns private writable signals** (source of truth) 2. **Exposes readonly signals to prevent direct mutation** 3. **Provides signal indices for efficient queries** (cached computed signals) 4. **Only allows changes via action dispatch** (predictable mutations) 5. **Actions automatically sync to realm peers** (P2P sync) 6. **Components read signals and dispatch actions** (reactive + predictable) This gives you: - **Redux-style predictable state management** - **Database-like query performance via reactive indices** - **Preact Signals efficient reactivity** - **Seamless P2P sync via your existing realm infrastructure** - **No boilerplate or manual subscriptions** Perfect local-first architecture that scales from simple to complex use cases. * uncategorized notes ** sync - each client keeps the full data set - dexie sync and observable let us stream change sets - we can publish the "latest" to all peers - on first pull, if not the first client, we can request a dump out of band *** rss feed data - do we want to backup feed data? - conceptually, this should be refetchable - but feeds go away, and some will only show recent stories - so yes, we'll need this - but server side, we can dedupe - content-addressed server-side cache? - server side does RSS pulling - can feeds be marked private, such that they won't be pulled through the proxy? - but then we require everything to be fetchable via cors - client configured proxy settings? *** peer connection - on startup, check for current realm-id and key pair - if not present, ask to login or start new - if login, run through the [[* pairing]] process - if start new, run through the [[* registration]] process - use keypair to authenticate to server - response includes list of active peers to connect - clients negotiate sync from there - an identity is a keypair and a realm - realm is uuid - realm on the server is the socket connection for peer discovery - keeps a list of verified public keys - and manages the /current/ ~public-key->peer ids~ mapping - realm on the client side is first piece of info required for sync - when connecting to the signalling server, you present a realm, and a signed public key - server accepts/rejects based on signature and current verified keys - a new keypair can create a realm - a new keypair can double sign an invitation - invite = ~{ realm:, nonce:, not_before:, not_after:, authorizer: }~, signed with verified key - exchanging an invite = ~{ invite: }~, signed with my key - on startup - start stand-alone (no syncing required, usually the case on first-run) - generate a keypair - want server backup? - sign a "setup" message with new keypair and send to the server - server responds with a new realm, that this keypair is already verified for - move along - exchange invite to sync to other devices - generate a keypair - sign the exchange message with the invite and send to the server - server verifies the invite - adds the new public key to the peer list and publishes downstream - move along ***** standalone in this mode, there is no syncing. this is the most likely first-time run option. - generate a keypair on startup, so we have a stable fingerprint in the future - done ***** pairing in this mode, there is syncing to a named realm, but not necessarily server resources consumed we don't need an email, since the server is just doing signalling and peer management - generate an invite from an existing verified peer - ~{ realm:, not_before:, not_after:, inviter: peer.public_key }~ - sign that invitation from the existing verified peer - standalone -> paired - get the invitation somehow (QR code?) - sign an invite exchange with the standalone's public key - send to server - server verifies the invite - adds the new public key to the peer list and publishes downstream ***** server backup in this mode, there is syncing to a named realm by email. goal of server backup mode is that we can go from email->fully working client with latest data without having to have any clients left around that could participate in the sync. - generate a keypair on startup - sign a registration message sent to the server - send a verification email - if email/realm already exists, this is authorization - if not, it's email validation - server starts a realm and associates the public key - server acts as a peer for the realm, and stores private data - since dexie is publishing change sets, we should be able to just store deltas - but we'll need to store _all_ deltas, unless we're materializing on the server side too - should we use an indexdb shim so we can import/export from the server for clean start? - how much materialization does the server need? ** summarized architecture design (may 28-29) :ai:claude: key decisions and system design: *** sync model - device-specific records for playback state/queues to avoid conflicts - content-addressed server cache with deduplication - dual-JWT invitation flow for secure realm joining *** data structures - tag-based filtering system instead of rigid hierarchies - regex patterns for episode title parsing and organization - service worker caching with background download support *** core schemas **** client (dexie) - Channel/ChannelEntry for RSS feeds and episodes - PlayRecord/QueueItem scoped by deviceId - FilterView for virtual feed organization **** server (drizzle) - ContentStore for deduplicated content by hash - Realm/PeerConnection for sync authorization - HttpCache with health tracking and TTL *** push sync strategy - revision-based sync (just send revision ranges in push notifications) - background fetch API for large downloads where supported - graceful degradation to reactive caching *** research todos :ai:claude: **** sync and data management ***** DONE identity and signature management ***** TODO dexie sync capabilities vs rxdb for multi-device sync implementation ***** TODO webrtc p2p sync implementation patterns and reliability ***** TODO conflict resolution strategies for device-specific data in distributed sync ***** TODO content-addressed deduplication algorithms for rss/podcast content **** client-side storage and caching ***** TODO opfs storage limits and cleanup strategies for client-side caching ***** TODO practical background fetch api limits and edge cases for podcast downloads **** automation and intelligence ***** TODO llm-based regex generation for episode title parsing automation ***** TODO push notification subscription management and realm authentication **** platform and browser capabilities ***** TODO browser audio api capabilities for podcast-specific features (speed, silence skip) ***** TODO progressive web app installation and platform-specific behaviors ** <2025-05-28 Wed> getting everything setup the biggest open question I have is what sort of privacy/encryption guarantee I need. I want the server to be able to do things like cache and store feed data long-term. Is "if you want full privacy, self-host" valid? *** possibilities - fully PWA - CON: cors, which would require a proxy anyway - CON: audio analysis, llm based stuff for categorization, etc. won't work - PRO: private as all get out - can still do WebRTC p2p sync for resiliancy - can still do server backups, if sync stream is encrypted, but no compaction would be available - could do _explicit_ server backups as dump files - self hostable - PRO: can do bunches of private stuff on the server, because if you don't want me to see it, do it elsewhere - CON: hard for folk to use *** sync conflict resolution design discussion :ai:claude: discussed the sync architecture and dexie conflict handling: *dexie syncable limitations*: - logical clocks handle causally-related changes well - basic timestamp-based conflict resolution for concurrent updates - last-writer-wins for same field conflicts - no sophisticated CRDT or vector clock support *solutions for podcast-specific conflicts*: - play records: device-specific approach - store separate ~play_records~ per ~device_id~ - each record: ~{ episode_id, device_id, position, completed, timestamp }~ - UI handles conflict resolution with "continue from X device?" prompts - avoids arbitrary timestamp wins, gives users control - subscription trees - store ~parent_path~ as single string field ("/Tech/Programming") - simpler than managing folder membership tables - conflicts still possible but contained to single field - could store move operations as events for richer resolution *other sync considerations*: - settings/preferences: distinguish device-local vs global - bulk operations: "mark all played" can create duplicate operations - metadata updates: server RSS updates vs local renames - temporal ordering: recently played lists, queue reordering - storage limits: cleanup operations conflicting across devices - feed state: refresh timestamps, error states *approach*: prefer "events not state" pattern and device-specific records where semantic conflicts are likely *** data model brainstorm :ai:claude: core entities designed with sync in mind: **** ~Feed~ :: RSS/podcast subscription - ~parent_path~ field for folder structure (eg. ~/Tech/Programming~) - ~is_private~ flag to skip server proxy - ~refresh_interval~ for custom update frequencies **** ~Episode~ :: individual podcast episodes - standard RSS metadata (guid, title, description, media url) - duration and file info for playback **** ~PlayRecord~ :: device-specific playback state - separate record per ~device_id~ to avoid timestamp conflicts - position, completed status, playback speed - UI can prompt "continue from X device?" for resolution **** ~QueueItem~ :: device-specific episode queue - ordered list with position field - ~device_id~ scoped to avoid queue conflicts **** ~Subscription~ :: feed membership settings - can be global or device-specific - auto-download preferences per device **** ~Settings~ :: split global vs device-local - theme, default speed = global - download path, audio device = device-local **** Event tables for complex operations: - ~FeedMoveEvent~ for folder reorganization - ~BulkMarkPlayedEvent~ for "mark all read" operations - better conflict resolution than direct state updates **** sync considerations - device identity established on first run - dexie syncable handles basic timestamp conflicts - prefer device-scoped records for semantic conflicts - event-driven pattern for bulk operations *** schema evolution from previous iteration :ai:claude: reviewed existing schema from tmp/feed.ts - well designed foundation: **** keep from original - Channel/ChannelEntry naming and structure - ~refreshHP~ adaptive refresh system (much better than simple intervals) - rich podcast metadata (people, tags, enclosure, podcast object) - HTTP caching with etag/status tracking - epoch millisecond timestamps - ~hashId()~ approach for entry IDs **** add for multi-device sync - ~PlayState~ table (device-scoped position/completion) - Subscription table (with ~parentPath~ for folders, device-scoped settings) - ~QueueItem~ table (device-scoped episode queues) - Device table (identity management) **** migration considerations - existing Channel/ChannelEntry can be preserved - new tables are additive - ~fetchAndUpsert~ method works well with server proxy architecture - dexie sync vs rxdb - need to evaluate change tracking capabilities *** content-addressed caching for offline resilience :ai:claude: designed caching system for when upstream feeds fail/disappear, building on existing cache-schema.ts: **** server-side schema evolution (drizzle sqlite): - keep existing ~httpCacheTable~ design (health tracking, http headers, ttl) - add ~contentHash~ field pointing to deduplicated content - new ~contentStoreTable~: deduplicated blobs by sha256 hash - new ~contentHistoryTable~: url -> contentHash timeline with isLatest flag - reference counting for garbage collection **** client-side OPFS storage - ~/cache/content/{contentHash}.xml~ for raw feeds - ~/cache/media/{contentHash}.mp3~ for podcast episodes - ~LocalCacheEntry~ metadata tracks expiration and offline-only flags - maintains last N versions per feed for historical access **** fetch strategy & fallback 1. check local OPFS cache first (fastest) 2. try server proxy ~/api/feed?url={feedUrl}~ (deduplicated) 3. server checks ~contentHistory~, serves latest or fetches upstream 4. server returns ~{contentHash, content, cached: boolean}~ 5. client stores with content hash as filename 6. emergency mode: serve stale content when upstream fails - preserves existing health tracking and HTTP caching logic - popular feeds cached once on server, many clients benefit - bandwidth savings via content hash comparison - historical feed state preservation (feeds disappear!) - true offline operation after initial sync ** <2025-05-29 Thu> :ai:claude: e2e encryption and invitation flow design worked through the crypto and invitation architecture. key decisions: *** keypair strategy - use jwk format for interoperability (server stores public keys) - ed25519 for signing, separate x25519 for encryption if needed - zustand lazy initialization pattern: ~ensureKeypair()~ on first use - store private jwk in persisted zustand state *** invitation flow: dual-jwt approach solved the chicken-and-egg problem of sharing encryption keys securely. **** qr code contains two signed jwts: 1. invitation token: ~{iss: inviter_fingerprint, sub: invitation_id, purpose: "realm_invite"}~ 2. encryption key token: ~{iss: inviter_fingerprint, ephemeral_private: base64_key, purpose: "ephemeral_key"}~ **** exchange process: 1. invitee posts jwt1 + their public keys to ~/invitations~ 2. server verifies jwt1 signature against realm members 3. if valid: adds invitee to realm, returns ~{realm_id, realm_members, encrypted_realm_key}~ 4. invitee verifies jwt2 signature against returned realm members 5. invitee extracts ephemeral private key, decrypts realm encryption key **** security properties: - server never has decryption capability (missing ephemeral private key) - both jwts must be signed by verified realm member - if first exchange fails, second jwt is cryptographically worthless - atomic operation: identity added only if invitation valid - built-in expiration and tamper detection via jwt standard **** considered alternatives: - raw ephemeral keys in qr: simpler but no authenticity - ecdh key agreement: chicken-and-egg problem with public key exchange - server escrow: good but missing authentication layer - password-based: requires secure out-of-band sharing the dual-jwt approach provides proper authenticated invitations while maintaining e2e encryption properties. **** refined dual-jwt with ephemeral signing simplified the approach by using ephemeral key for second jwt signature: **setup**: 1. inviter generates ephemeral keypair 2. encrypts realm key with ephemeral private key 3. posts to server: ~{invitation_id, realm_id, ephemeral_public, encrypted_realm_key}~ **qr code contains**: #+BEGIN_SRC json // JWT 1: signed with inviter's realm signing key { "realm_id": "uuid", "invitation_id": "uuid", "iss": "inviter_fingerprint" } // JWT 2: signed with ephemeral private key { "ephemeral_private": "base64_key", "invitation_id": "uuid" } #+END_SRC **exchange flow**: 1. submit jwt1 → server verifies against realm members → returns ~{invitation_id, realm_id, ephemeral_public, encrypted_realm_key}~ 2. verify jwt2 signature using ~ephemeral_public~ from server response 3. extract ~ephemeral_private~ from jwt2, decrypt realm key **benefits over previous version**: - no premature key disclosure (invitee keys shared via normal webrtc peering) - self-contained verification (ephemeral public key verifies jwt2) - cleaner separation of realm auth vs encryption key distribution - simpler flow (no need to return realm member list) **crypto verification principle**: digital signatures work as sign-with-private/verify-with-public, while encryption works as encrypt-with-public/decrypt-with-private. jwt2 verification uses signature verification, not decryption. **invitation flow diagram**: #+BEGIN_SRC mermaid sequenceDiagram participant I as Inviter participant S as Server participant E as Invitee Note over I: Generate ephemeral keypair I->>I: ephemeral_private, ephemeral_public Note over I: Encrypt realm key I->>I: encrypted_realm_key = encrypt(realm_key, ephemeral_private) I->>S: POST /invitations
{invitation_id, realm_id, ephemeral_public, encrypted_realm_key} S-->>I: OK Note over I: Create JWTs for QR code I->>I: jwt1 = sign({realm_id, invitation_id}, inviter_private) I->>I: jwt2 = sign({ephemeral_private, invitation_id}, ephemeral_private) Note over I,E: QR code contains [jwt1, jwt2] E->>S: POST /invitations/exchange
{jwt1} Note over S: Verify jwt1 signature
against realm members S-->>E: {invitation_id, realm_id, ephemeral_public, encrypted_realm_key} Note over E: Verify jwt2 signature
using ephemeral_public E->>E: verify_signature(jwt2, ephemeral_public) Note over E: Extract key and decrypt E->>E: ephemeral_private = decode(jwt2) E->>E: realm_key = decrypt(encrypted_realm_key, ephemeral_private) Note over E: Now member of realm! #+END_SRC **** jwk keypair generation and validation :ai:claude: discussed jwk vs raw crypto.subtle for keypair storage. since public keys need server storage for realm authorization, jwk is better for interoperability. **keypair generation**: #+BEGIN_SRC typescript const keypair = await crypto.subtle.generateKey( { name: "Ed25519" }, true, ["sign", "verify"] ); const publicJWK = await crypto.subtle.exportKey("jwk", keypair.publicKey); const privateJWK = await crypto.subtle.exportKey("jwk", keypair.privateKey); // JWK format: { "kty": "OKP", "crv": "Ed25519", "x": "base64url-encoded-public-key", "d": "base64url-encoded-private-key" // only in private JWK } #+END_SRC **client validation**: #+BEGIN_SRC typescript function isValidEd25519PublicJWK(jwk: any): boolean { return ( typeof jwk === 'object' && jwk.kty === 'OKP' && jwk.crv === 'Ed25519' && typeof jwk.x === 'string' && jwk.x.length === 43 && // base64url Ed25519 public key length !jwk.d && // public key shouldn't have private component !jwk.use || jwk.use === 'sig' ); } async function validatePublicKey(publicJWK: JsonWebKey): Promise { try { if (!isValidEd25519PublicJWK(publicJWK)) return null; const key = await crypto.subtle.importKey( 'jwk', publicJWK, { name: 'Ed25519' }, false, ['verify'] ); return key; } catch { return null; } } #+END_SRC **server validation (node.js)**: #+BEGIN_SRC typescript import { webcrypto } from 'node:crypto'; async function validateClientPublicKey(publicJWK: JsonWebKey): Promise { try { if (!isValidEd25519PublicJWK(publicJWK)) return false; await webcrypto.subtle.importKey( 'jwk', publicJWK, { name: 'Ed25519' }, false, ['verify'] ); return true; } catch { return false; } } #+END_SRC **authentication flow**: #+BEGIN_SRC typescript // client signs message const authMessage = { realm: 'uuid-here', timestamp: Date.now(), action: 'join' }; const signature = await crypto.subtle.sign( 'Ed25519', privateKey, new TextEncoder().encode(JSON.stringify(authMessage)) ); // server verifies async function verifyAuth(req: AuthRequest): Promise { const publicKey = await webcrypto.subtle.importKey( 'jwk', req.publicKey, { name: 'Ed25519' }, false, ['verify'] ); const messageBytes = new TextEncoder().encode(JSON.stringify(req.message)); const signatureBytes = new Uint8Array(req.signature); return await webcrypto.subtle.verify( 'Ed25519', publicKey, signatureBytes, messageBytes ); } #+END_SRC **** proposed schemas :ai:claude: ***** client-side schema (dexie) #+BEGIN_SRC typescript // Core RSS/Podcast data (from your existing design) interface Channel { id: string; feedUrl: string; htmlUrl?: string; imageUrl?: string; title?: string; description?: string; language?: string; people?: Record; tags?: string[]; // Refresh management refreshHP: number; nextRefreshAt?: number; lastRefreshAt?: number; lastRefreshStatus?: string; lastRefreshHttpStatus?: number; lastRefreshHttpEtag?: string; // Cache info contentHash?: string; lastFetchedAt?: number; } interface ChannelEntry { id: string; channelId: string; guid: string; title: string; linkUrl?: string; imageUrl?: string; snippet?: string; content?: string; enclosure?: { url: string; type?: string; length?: number; }; podcast?: { explicit?: boolean; duration?: string; seasonNum?: number; episodeNum?: number; transcriptUrl?: string; }; publishedAt?: number; fetchedAt?: number; } // Device-specific sync tables interface PlayRecord { id: string; entryId: string; deviceId: string; position: number; duration?: number; completed: boolean; speed: number; updatedAt: number; } interface Subscription { id: string; channelId: string; deviceId?: string; parentPath: string; // "/Tech/Programming" autoDownload: boolean; downloadLimit?: number; isActive: boolean; createdAt: number; updatedAt: number; } interface QueueItem { id: string; entryId: string; deviceId: string; position: number; addedAt: number; } interface Device { id: string; name: string; platform: string; lastSeen: number; } // Local cache metadata interface LocalCache { id: string; url: string; contentHash: string; filePath: string; // OPFS path cachedAt: number; expiresAt?: number; size: number; isOfflineOnly: boolean; } // Dexie schema const db = new Dexie('SkypodDB'); db.version(1).stores({ channels: '&id, feedUrl, contentHash', channelEntries: '&id, channelId, publishedAt', playRecords: '&id, [entryId+deviceId], deviceId, updatedAt', subscriptions: '&id, channelId, deviceId, parentPath', queueItems: '&id, entryId, deviceId, position', devices: '&id, lastSeen', localCache: '&id, url, contentHash, expiresAt' }); #+END_SRC ***** server-side schema #+BEGIN_SRC typescript // Content-addressed cache interface ContentStore { contentHash: string; // Primary key content: Buffer; // Raw feed content contentType: string; contentLength: number; firstSeenAt: number; referenceCount: number; } interface ContentHistory { id: string; url: string; contentHash: string; fetchedAt: number; isLatest: boolean; } // HTTP cache with health tracking (from your existing design) interface HttpCache { key: string; // URL hash, primary key url: string; status: 'alive' | 'dead'; lastFetchedAt: number; lastFetchError?: string; lastFetchErrorStreak: number; lastHttpStatus: number; lastHttpEtag?: string; lastHttpHeaders: Record; expiresAt: number; expirationTtl: number; contentHash: string; // Points to ContentStore } // Sync/auth tables interface Realm { id: string; // UUID createdAt: number; verifiedKeys: string[]; // Public key list } interface PeerConnection { id: string; realmId: string; publicKey: string; lastSeen: number; isOnline: boolean; } // Media cache for podcast episodes interface MediaCache { contentHash: string; // Primary key originalUrl: string; mimeType: string; fileSize: number; content: Buffer; cachedAt: number; accessCount: number; } #+END_SRC **** episode title parsing for sub-feed groupings :ai:claude: *problem*: some podcast feeds contain multiple shows, need hierarchical organization within a feed *example*: "Apocalypse Players" podcast - episode title: "A Term of Art 6 - Winston's Hollow" - desired grouping: "Apocalypse Players > A Term of Art > 6 - Winston's Hollow" - UI shows sub-shows within the main feed ***** approaches considered 1. *manual regex patterns* (short-term solution) - user provides regex with capture groups = tags - reliable, immediate, user-controlled - requires manual setup per feed 2. *LLM-generated regex* (automation goal) - analyze last 100 episode titles - generate regex pattern automatically - good balance of automation + reliability 3. *NER model training* (experimental) - train spacy model for episode title parsing - current prototype: 150 labelled examples, limited success - needs more training data to be viable ***** data model implications - add regex pattern field to Channel/Feed - store extracted groupings as hierarchical tags on ~ChannelEntry~ - maybe add grouping/series field to episodes ***** plan *preference*: start with manual regex, evolve toward LLM automation *implementation design*: - if no title pattern: episodes are direct children of the feed - title pattern = regex with named capture groups + path template *example configuration*: - regex: ~^(?[^0-9]+)\s*(?\d+)\s*-\s*(?.+)$~ - path template: ~{series} > Episode {episode} - {title}~ - result: "A Term of Art 6 - Winston's Hollow" → "A Term of Art > Episode 6 - Winston's Hollow" *schema additions*: #+BEGIN_SRC typescript interface Channel { // ... existing fields titlePatterns?: Array<{ name: string; // "Main Episodes", "Bonus Content", etc. regex: string; // named capture groups pathTemplate: string; // interpolation template priority: number; // order to try patterns (lower = first) isActive: boolean; // can disable without deleting }>; fallbackPath?: string; // template for unmatched episodes } interface ChannelEntry { // ... existing fields parsedPath?: string; // computed from titlePattern parsedGroups?: Record<string, string>; // captured groups matchedPatternName?: string; // which pattern was used } #+END_SRC *pattern matching logic*: 1. try patterns in priority order (lower number = higher priority) 2. first matching pattern wins 3. if no patterns match, use fallbackPath template (e.g., "Misc > {title}") 4. if no fallbackPath, episode stays direct child of feed *example multi-pattern setup*: - Pattern 1: "Main Episodes" - ~^(?<series>[^0-9]+)\s*(?<episode>\d+)~ → ~{series} > Episode {episode}~ - Pattern 2: "Bonus Content" - ~^Bonus:\s*(?<title>.+)~ → ~Bonus > {title}~ - Fallback: ~Misc > {title}~ **** scoped tags and filter-based UI evolution :ai:claude: *generalization*: move from rigid hierarchies to tag-based filtering system *tag scoping*: - feed-level tags: "Tech", "Gaming", "D&D" - episode-level tags: from regex captures like "series:CriticalRole", "campaign:2", "type:main" - user tags: manual additions like "favorites", "todo" *UI as tag filtering*: - default view: all episodes grouped by feed - filter by ~series:CriticalRole~ → shows only CR episodes across all feeds - filter by ~type:bonus~ → shows bonus content from all podcasts - combine filters: ~series:CriticalRole AND type:main~ → main CR episodes only *benefits*: - no rigid hierarchy - users create their own views - regex patterns become automated episode taggers - same filtering system works for search, organization, queues - tags are syncable metadata, views are client-side *schema evolution*: #+BEGIN_SRC typescript interface Tag { scope: 'feed' | 'episode' | 'user'; key: string; // "series", "type", "campaign" value: string; // "CriticalRole", "bonus", "2" } interface ChannelEntry { // ... existing tags: Tag[]; // includes regex-generated + manual } interface FilterView { id: string; name: string; folderPath: string; // "/Channels/Critical Role" filters: Array<{ key: string; value: string; operator: 'equals' | 'contains' | 'not'; }>; isDefault: boolean; createdAt: number; } #+END_SRC **** default UI construction and feed merging :ai:claude: *auto-generated views on subscribe*: - subscribe to "Critical Role" → creates ~/Channels/Critical Role~ folder - default filter view: ~feed:CriticalRole~ (shows all episodes from that feed) - user can customize, split into sub-views, or delete *smart view suggestions*: - after regex patterns generate tags, suggest splitting views - "I noticed episodes with ~series:Campaign2~ and ~series:Campaign3~ - create separate views?" - "Create view for ~type:bonus~ episodes?" *view management UX*: - right-click feed → "Split by series", "Split by type" - drag episodes between views to create manual filters - views can be nested: ~/Channels/Critical Role/Campaign 2/Main Episodes~ *feed merging for multi-source shows*: problem: patreon feed + main show feed for same podcast #+BEGIN_EXAMPLE /Channels/ Critical Role/ All Episodes # merged view: feed:CriticalRole OR feed:CriticalRolePatreon Main Feed # filter: feed:CriticalRole Patreon Feed # filter: feed:CriticalRolePatreon #+END_EXAMPLE *deduplication strategy*: - episodes matched by ~guid~ or similar content hash - duplicate episodes get ~source:main,patreon~ tags - UI shows single episode with source indicators - user can choose preferred source for playback - play state syncs across all sources of same episode *feed relationship schema*: #+BEGIN_SRC typescript interface FeedGroup { id: string; name: string; // "Critical Role" feedIds: string[]; // [mainFeedId, patreonFeedId] mergeStrategy: 'guid' | 'title' | 'contentHash'; defaultView: FilterView; } interface ChannelEntry { // ... existing duplicateOf?: string; // points to canonical episode ID sources: string[]; // feed IDs where this episode appears } #+END_SRC **per-view settings and state**: each filter view acts like a virtual feed with its own: - unread counts (episodes matching filter that haven't been played) - notification settings (notify for new episodes in this view) - muted state (hide notifications, mark as read automatically) - auto-download preferences (download episodes that match this filter) - play queue integration (add new episodes to queue) **use cases**: - mute "Bonus Content" view but keep notifications for main episodes - auto-download only "Campaign 2" episodes, skip everything else - separate unread counts: "5 unread in Main Episodes, 2 in Bonus" - queue only certain series automatically **schema additions**: #+BEGIN_SRC typescript interface FilterView { // ... existing fields settings: { notificationsEnabled: boolean; isMuted: boolean; autoDownload: boolean; autoQueue: boolean; downloadLimit?: number; // max episodes to keep }; state: { unreadCount: number; lastViewedAt?: number; isCollapsed: boolean; // in sidebar }; } #+END_SRC *inheritance behavior*: - new filter views inherit settings from parent feed/group - user can override per-view - "mute all Critical Role" vs "mute only bonus episodes" **** client-side episode caching strategy :ai:claude: *architecture*: service worker-based transparent caching *flow*: 1. audio player requests ~/audio?url={episodeUrl}~ 2. service worker intercepts request 3. if present in cache (with Range header support): - serve from cache 4. else: - let request continue to server (immediate playback) - simultaneously start background fetch of full audio file - when complete, broadcast "episode-cached" event - audio player catches event and restarts feed → now uses cached version **benefits**: - no playback interruption (streaming starts immediately) - seamless transition to cached version - Range header support for seeking/scrubbing - transparent to audio player implementation *implementation considerations*: - cache storage limits and cleanup policies - partial download resumption if interrupted - cache invalidation when episode URLs change - offline playback support - progress tracking for background downloads **schema additions**: #+BEGIN_SRC typescript interface CachedEpisode { episodeId: string; originalUrl: string; cacheKey: string; // for cache API fileSize: number; cachedAt: number; lastAccessedAt: number; downloadProgress?: number; // 0-100 for in-progress downloads } #+END_SRC **service worker events**: - ~episode-cache-started~ - background download began - ~episode-cache-progress~ - download progress update - ~episode-cache-complete~ - ready to switch to cached version - ~episode-cache-error~ - download failed, stay with streaming **background sync for proactive downloads**: **browser support reality**: - Background Sync API: good support (Chrome/Edge, limited Safari) - Periodic Background Sync: very limited (Chrome only, requires PWA install) - Push notifications: good support, but requires user permission **hybrid approach**: 1. **foreground sync** (reliable): when app is open, check for new episodes 2. **background sync** (opportunistic): register sync event when app closes 3. **push notifications** (fallback): server pushes "new episodes available" 4. **manual sync** (always works): pull-to-refresh, settings toggle **implementation strategy**: #+BEGIN_SRC typescript // Register background sync when app becomes hidden document.addEventListener('visibilitychange', () => { if (document.hidden && 'serviceWorker' in navigator) { navigator.serviceWorker.ready.then(registration => { return registration.sync.register('download-episodes'); }); } }); // Service worker handles sync event self.addEventListener('sync', event => { if (event.tag === 'download-episodes') { event.waitUntil(syncEpisodes()); } }); #+END_SRC **realistic expectations**: - iOS Safari: very limited background processing - Android Chrome: decent background sync support - Desktop: mostly works - battery/data saver modes: disabled by OS **fallback strategy**: rely primarily on foreground sync + push notifications, treat background sync as nice-to-have enhancement **push notification sync workflow**: **server-side trigger**: 1. server detects new episodes during RSS refresh 2. check which users are subscribed to that feed 3. send push notification with episode metadata payload 4. notification wakes up service worker on client **service worker notification handler**: #+BEGIN_SRC typescript self.addEventListener('push', event => { const data = event.data?.json(); if (data.type === 'new-episodes') { event.waitUntil( // Start background download of new episodes downloadNewEpisodes(data.episodes) .then(() => { // Show notification to user return self.registration.showNotification('New episodes available', { body: ~${data.episodes.length} new episodes downloaded~, icon: '/icon-192.png', badge: '/badge-72.png', tag: 'new-episodes', data: { episodeIds: data.episodes.map(e => e.id) } }); }) ); } }); // Handle notification click self.addEventListener('notificationclick', event => { event.notification.close(); // Open app to specific episode or feed event.waitUntil( clients.openWindow(~/episodes/${event.notification.data.episodeIds[0]}~) ); }); #+END_SRC **server push logic**: - batch notifications (don't spam for every episode) - respect user notification preferences from FilterView settings - include episode metadata in payload to avoid round-trip - throttle notifications (max 1 per feed per hour?) **user flow**: 1. new episode published → server pushes notification 2. service worker downloads episode in background 3. user sees "New episodes downloaded" notification 4. tap notification → opens app to new episode, ready to play offline *benefits*: - true background downloading without user interaction - works even when app is closed - respects per-feed notification settings **push payload size constraints**: - **limit**: ~4KB (4,096 bytes) across most services - **practical limit**: ~3KB to account for service overhead - **implications for episode metadata**: #+BEGIN_SRC json { "type": "new-episodes", "episodes": [ { "id": "ep123", "channelId": "ch456", "title": "Episode Title", "url": "https://...", "duration": 3600, "size": 89432112 } ] } #+END_SRC **payload optimization strategies**: - minimal episode metadata in push (id, url, basic info) - batch multiple episodes in single notification - full episode details fetched after service worker wakes up - URL shortening for long episode URLs - compress JSON payload if needed **alternative for large payloads**: - push notification contains only "new episodes available" signal - service worker makes API call to get full episode list - trade-off: requires network round-trip but unlimited data **logical clock sync optimization**: much simpler approach using sync revisions: #+BEGIN_SRC json { "type": "sync-available", "fromRevision": 12345, "toRevision": 12389, "changeCount": 8 } #+END_SRC **service worker sync flow**: 1. push notification wakes service worker with revision range 2. service worker fetches ~/sync?from=12345&to=12389~ 3. server returns only changes in that range (episodes, feed updates, etc) 4. service worker applies changes to local dexie store 5. service worker queues background downloads for new episodes 6. updates local revision to 12389 **benefits of revision-based approach**: - tiny push payload (just revision numbers) - server can efficiently return only changes in range - automatic deduplication (revision already applied = skip) - works for any sync data (episodes, feed metadata, user settings) - handles offline gaps gracefully (fetch missing revision ranges) **sync API response**: #+BEGIN_SRC typescript interface SyncResponse { fromRevision: number; toRevision: number; changes: Array<{ type: 'episode' | 'channel' | 'subscription'; operation: 'create' | 'update' | 'delete'; data: any; revision: number; }>; } #+END_SRC **integration with episode downloads**: - service worker processes sync changes - identifies new episodes that match user's auto-download filters - queues those for background cache fetching - much more efficient than sending episode metadata in push payload **service worker processing time constraints**: **hard limits**: - **30 seconds idle timeout**: service worker terminates after 30s of inactivity - **5 minutes event processing**: single event/request must complete within 5 minutes - **30 seconds fetch timeout**: individual network requests timeout after 30s - **notification requirement**: push events MUST display notification before promise settles **practical implications**: - sync API call (~/sync?from=X&to=Y~) must complete within 30s - large episode downloads must be queued, not started immediately in push handler - use ~event.waitUntil()~ to keep service worker alive during processing - break large operations into smaller chunks **recommended push event flow**: #+BEGIN_SRC typescript self.addEventListener('push', event => { const data = event.data?.json(); event.waitUntil( // Must complete within 5 minutes total handlePushSync(data) .then(() => { // Required: show notification before promise settles return self.registration.showNotification('Episodes synced'); }) ); }); async function handlePushSync(data) { // 1. Quick sync API call (< 30s) const changes = await fetch(~/sync?from=${data.fromRevision}&to=${data.toRevision}~); // 2. Apply changes to dexie store (fast, local) await applyChangesToStore(changes); // 3. Queue episode downloads for later (don't start here) await queueEpisodeDownloads(changes.newEpisodes); // Total time: < 5 minutes, preferably < 30s } #+END_SRC *download strategy*: use push event for sync + queuing, separate background tasks for actual downloads *background fetch API for large downloads*: *progressive enhancement approach*: #+BEGIN_SRC typescript async function queueEpisodeDownloads(episodes) { for (const episode of episodes) { if ('serviceWorker' in navigator && 'BackgroundFetch' in window) { // Chrome/Edge: use Background Fetch API for true background downloading await navigator.serviceWorker.ready.then(registration => { return registration.backgroundFetch.fetch( ~episode-${episode.id}~, episode.url, { icons: [{ src: '/icon-256.png', sizes: '256x256', type: 'image/png' }], title: ~Downloading: ${episode.title}~, downloadTotal: episode.fileSize } ); }); } else { // Fallback: queue for reactive download (download while streaming) await queueReactiveDownload(episode); } } } // Handle background fetch completion self.addEventListener('backgroundfetch', event => { if (event.tag.startsWith('episode-')) { event.waitUntil(handleEpisodeDownloadComplete(event)); } }); #+END_SRC *browser support reality*: - *Chrome/Edge*: Background Fetch API supported - *Firefox/Safari*: not supported, fallback to reactive caching - *mobile*: varies by platform and browser *benefits when available*: - true background downloading (survives app close, browser close) - built-in download progress UI - automatic retry on network failure - no service worker time limits during download *graceful degradation*: - detect support, use when available - fallback to reactive caching (download while streaming) - user gets best experience possible on their platform *** research todos :ai:claude: high-level unanswered questions from architecture brainstorming: **** sync and data management ***** TODO dexie sync capabilities vs rxdb for multi-device sync implementation ***** TODO webrtc p2p sync implementation patterns and reliability ***** TODO conflict resolution strategies for device-specific data in distributed sync ***** TODO content-addressed deduplication algorithms for rss/podcast content **** client-side storage and caching ***** TODO opfs storage limits and cleanup strategies for client-side caching ***** TODO practical background fetch api limits and edge cases for podcast downloads **** automation and intelligence ***** TODO llm-based regex generation for episode title parsing automation ***** TODO push notification subscription management and realm authentication **** platform and browser capabilities ***** TODO browser audio api capabilities for podcast-specific features (speed, silence skip) ***** TODO progressive web app installation and platform-specific behaviors * webtorrent brainstorming 6/16 :ai:claude: ** WebTorrent + Event Log CRDT Architecture *** Core Concept Split We identified two fundamentally different types of data that need different sync strategies: **** 1. Dynamic Metadata (Event Log CRDT) - **Data**: Play state, scroll position, settings, subscriptions - **Characteristics**: Frequently changing, small, device-specific - **Solution**: Event log with Hybrid Logical Clocks (HLC) - **Sync**: Merkle tree efficient diff + P2P exchange via realm **** 2. Static Content (WebTorrent) - **Data**: RSS feeds, podcast episodes (audio files) - **Characteristics**: Immutable, large, content-addressable - **Solution**: WebTorrent with infohash references - **Storage**: IndexedDB chunk store (idb-chunk-store npm package) *** Event Log CRDT Design **** Hybrid Logical Clock (HLC) Based on James Long's crdt-example-app implementation: #+BEGIN_SRC typescript interface HLC { millis: number; // physical time counter: number; // logical counter (0-65535) node: string; // device identity ID } interface SyncEvent { timestamp: HLC; type: 'subscribe' | 'unsubscribe' | 'markPlayed' | 'updatePosition' | ... payload: any; } #+END_SRC **Benefits**: - Causality preserved even with clock drift - Compact representation (vs full vector clocks) - Total ordering via (millis, counter, node) comparison - No merge conflicts - just union of events **** Merkle Tree Sync Efficient sync using merkle trees over time ranges: #+BEGIN_SRC typescript interface RangeMerkleNode { startTime: HLC; endTime: HLC; hash: string; eventCount: number; } #+END_SRC **Sync Protocol**: 1. Exchange merkle roots 2. If different, drill down to find divergent ranges 3. Exchange only missing events 4. Apply in HLC order **Key insight**: No merge conflicts because events are immutable and ordered by HLC **** Progressive Compaction Use idle time to compact old events: - Recent (< 5 min): Individual events for active sync - Hourly chunks: After 5 minutes - Daily chunks: After 24 hours - Monthly chunks: After 30 days Benefits: - Fast recent sync - Efficient storage of history - Old chunks can move to OPFS as blobs *** WebTorrent Integration **** Content Flow 1. **CORS-friendly feeds**: - Browser fetches directly - Creates torrent with original URL as webseed - Broadcasts infohash to realm 2. **CORS-blocked feeds**: - Server fetches and hashes - Returns infohash (server doesn't store content) - Client uses WebTorrent with original URL as webseed **** Realm as Private Tracker - Realm members announce infohashes they have - No need for DHT or public trackers - Existing WebRTC signaling used for peer discovery - Private swarm for each realm **** Storage via Chunk Store Use `idb-chunk-store` (or similar) for persistence: - WebTorrent handles chunking/verification - IndexedDB provides persistence across sessions - Abstract-chunk-store interface allows swapping implementations *** Bootstrap & History Sharing **** History Snapshots as Torrents Serialize event history into content-addressed chunks: #+BEGIN_SRC typescript interface HistorySnapshot { period: "2024-05"; events: SyncEvent[]; merkleRoot: string; deviceStates: Record<string, DeviceState>; } // Share via WebTorrent const blob = await serializeSnapshot(events); const infohash = await createTorrent(blob); realm.broadcast({ type: "historySnapshot", period, infohash }); #+END_SRC **** Materialized State Snapshots Using dexie-export-import for database snapshots: #+BEGIN_SRC typescript const dbBlob = await exportDB(db, { tables: ['channels', 'channelEntries'], filter: (table, value) => !isDeviceSpecific(table, value) }); const infohash = await createTorrent(dbBlob); #+END_SRC **** New Device Bootstrap 1. Download latest DB snapshot → Instant UI 2. Download recent events → Apply updates 3. Background: fetch historical event logs 4. Result: Fast startup with complete history *** Implementation Benefits 1. **Privacy**: No server sees listening history 2. **Offline-first**: Everything works locally 3. **Efficient sync**: Only exchange missing data 4. **P2P content**: Reduce server bandwidth 5. **Scalable**: Torrents for bulk data transfer 6. **Verifiable**: Merkle trees ensure consistency *** Next Steps - [ ] Implement HLC timestamps - [ ] Build merkle tree sync protocol - [ ] Integrate WebTorrent with realm signaling - [ ] Create history snapshot system - [ ] Test cross-device sync scenarios ** Additional Architecture Insights *** Unified Infohash Approach Instead of having separate hashes for merkle tree and WebTorrent, use infohashes throughout: **** Hierarchical Infohash Structure #+BEGIN_SRC typescript // Leaf level: individual files const episode1Hash = await createTorrent(episode1.mp3); const feedXmlHash = await createTorrent(feed.xml); // Directory level: multi-file torrent const feedTorrent = await createTorrent({ name: 'example.com.rss', files: [ { path: 'rss.xml', infohash: feedXmlHash }, { path: 'episode-1.mp3', infohash: episode1Hash } ] }); // Root level: torrent of feed torrents const rootTorrent = await createTorrent({ name: 'feeds', folders: [ { path: 'example.com.rss', infohash: feedTorrent.infoHash } ] }); #+END_SRC Benefits: - Single hash type throughout system - Progressive loading (directory structure first, then files) - Natural deduplication - WebTorrent native sharing of folder structures *** Long-term Event Log Scaling **** Checkpoint + Delta Pattern For handling millions of events, use periodic checkpoints: #+BEGIN_SRC typescript interface EventCheckpoint { hlc: HLC; stateSnapshot: { subscriptions: Channel[]; playStates: PlayRecord[]; settings: Settings; }; eventCount: number; infohash: string; // torrent of this checkpoint } // Every 10k events or monthly async function createCheckpoint(): Promise<Checkpoint> { const currentHLC = getLatestEventHLC(); // Export materialized state using dexie-export-import const dbBlob = await exportDB(db, { filter: (table, value) => { return !['activeSyncs', 'tempData'].includes(table); } }); const infohash = await createTorrent(dbBlob); return { hlc: currentHLC, dbExport: dbBlob, infohash }; } #+END_SRC **** Bootstrap Flow with Checkpoints 1. New device downloads latest checkpoint via WebTorrent 2. Imports directly to IndexedDB: `await importDB(checkpoint.blob)` 3. Requests only recent events since checkpoint 4. Applies recent events to catch up Benefits: - Fast bootstrap (one checkpoint instead of million events) - No double materialization (IndexedDB is already materialized state) - P2P distribution of checkpoints - Clear version migration path *** Sync State Management **** Catching Up vs Live Events #+BEGIN_SRC typescript interface SyncState { localHLC: HLC; remoteHLC: HLC; mode: 'catching-up' | 'live'; } // Separate handlers for historical vs live events async function replayHistoricalEvents(from: HLC, to: HLC) { const events = await fetchEvents(from, to); // Process in batches without UI updates await db.transaction('rw', db.tables, async () => { for (const batch of chunks(events, 1000)) { await Promise.all(batch.map(applyEventSilently)); } }); // One UI update at the end notifyUI('Sync complete', { newEpisodes: 47 }); } function handleLiveEvent(event: SyncEvent) { // Real-time event - update UI immediately applyEvent(event); if (event.type === 'newEpisode') { showNotification(`New episode: ${event.title}`); } } #+END_SRC **** HLC Comparison for Ordering #+BEGIN_SRC typescript function compareHLC(a: HLC, b: HLC): number { if (a.millis !== b.millis) return a.millis - b.millis; if (a.counter !== b.counter) return a.counter - b.counter; return a.node.localeCompare(b.node); } // Determine if caught up function isCaughtUp(myHLC: HLC, peerHLC: HLC): boolean { return compareHLC(myHLC, peerHLC) >= 0; } #+END_SRC *** Handling Out-of-Order Events **** Idempotent Reducers (No Replay Needed) Design reducers to handle events arriving out of order: #+BEGIN_SRC typescript // HLC-aware reducer that handles out-of-order events function reducePlayPosition(state, event) { if (event.type === 'updatePosition') { const existing = state.positions[event.episodeId]; // Only update if this event is newer if (!existing || compareHLC(event.hlc, existing.hlc) > 0) { state.positions[event.episodeId] = { position: event.position, hlc: event.hlc // Track which event set this }; } } } #+END_SRC **** Example: Offline Device Rejoining #+BEGIN_SRC typescript // Device A offline for a week, comes back with old events Device A: [ { hlc: "1000:0:A", type: "markPlayed", episode: "ep1" }, { hlc: "1100:0:A", type: "updatePosition", episode: "ep1", position: 500 } ] // Device B already has newer event Device B: [ { hlc: "1050:0:B", type: "updatePosition", episode: "ep1", position: 1000 } ] // Smart reducer produces correct final state finalState = { "ep1": { played: true, // from 1000:0:A position: 1000, // from 1050:0:B (newer HLC wins) lastPositionHLC: "1050:0:B" } } #+END_SRC Key principles: - Store HLC with state changes - Use "last write wins" with HLC comparison - Make operations commutative when possible - No need for full replay when inserting old events * Redux-style Action Layer Implementation Plan :ai:claude: ** Overview Implementation of Redux-style action/reducer pattern with Logical Clock timestamps for P2P sync via WebRTC. All state changes flow through actions, with Dexie used as the persistence layer (implementation detail of the reducer). ** Core Components *** 1. Base Action Schema (`src/common/protocol/actions.ts`) #+BEGIN_SRC typescript import {z} from 'zod/v4' import {LogicalClock} from './logical-clock' // Base action schema with logical clock timestamp export const actionSchema = z.object({ type: z.string(), payload: z.unknown(), timestamp: LogicalClock.schema.optional(), meta: z.object({ skipSync: z.boolean().optional(), }).optional(), }) // Helper to create typed action schemas export const makeActionSchema = <T extends string, P extends z.ZodType>( type: T, payload: P ) => { return actionSchema.extend({ type: z.literal(type), payload, }) } export type Action = z.infer<typeof actionSchema> #+END_SRC *** 2. Skypod Action Definitions (`src/common/protocol/skypod-actions.ts`) #+BEGIN_SRC typescript import {z} from 'zod/v4' import {makeActionSchema} from './actions' // Channel actions export const channelSubscribeActionSchema = makeActionSchema( 'channel:subscribe', z.object({ url: z.string().url(), title: z.string().optional(), }) ) export const channelUpdateActionSchema = makeActionSchema( 'channel:update', z.object({ guid: z.string(), title: z.string().optional(), description: z.string().optional(), imageUrl: z.string().url().optional(), }) ) export const channelUnsubscribeActionSchema = makeActionSchema( 'channel:unsubscribe', z.object({ guid: z.string(), }) ) // Entry actions export const entryMarkReadActionSchema = makeActionSchema( 'entry:markRead', z.object({ channelGuid: z.string(), entryGuid: z.string(), readAt: z.number(), }) ) // Union of all actions for validation export const skypodActionSchema = z.discriminatedUnion('type', [ channelSubscribeActionSchema, channelUpdateActionSchema, channelUnsubscribeActionSchema, entryMarkReadActionSchema, ]) // Export types export type ChannelSubscribeAction = z.infer<typeof channelSubscribeActionSchema> export type ChannelUpdateAction = z.infer<typeof channelUpdateActionSchema> export type ChannelUnsubscribeAction = z.infer<typeof channelUnsubscribeActionSchema> export type EntryMarkReadAction = z.infer<typeof entryMarkReadActionSchema> export type SkypodAction = z.infer<typeof skypodActionSchema> #+END_SRC *** 3. Sync Protocol Messages (`src/common/protocol/messages.ts`) Add to existing messages: #+BEGIN_SRC typescript // Sync request - peer requests actions since timestamp export const skypodSyncRequestSchema = makeRequestSchema( 'skypod.sync', z.object({ since: LogicalClock.schema.optional(), // undefined = get all }) ) // Sync response - peer sends requested actions export const skypodSyncResponseSchema = makeResponseSchema( 'skypod.sync', z.object({ actions: z.array(skypodActionSchema), }) ) // Sync event - peer broadcasts new action in real-time export const skypodSyncEventSchema = makeEventSchema( 'skypod.sync', skypodActionSchema ) export type SkypodSyncRequest = z.infer<typeof skypodSyncRequestSchema> export type SkypodSyncResponse = z.infer<typeof skypodSyncResponseSchema> export type SkypodSyncEvent = z.infer<typeof skypodSyncEventSchema> #+END_SRC *** 4. Database Schema Updates (`src/client/database.ts`) #+BEGIN_SRC typescript import {LCTimestamp} from '#common/protocol/logical-clock' import {SkypodAction} from '#common/protocol/skypod-actions' // Add timestamp to Channel for conflict resolution export interface Channel { // ... existing fields ... timestamp?: LCTimestamp } // New table for action history export interface StoredAction extends SkypodAction { timestamp: LCTimestamp // Required for stored actions } export class SkypodDatabase extends Dexie { // ... existing tables ... actions!: Table<StoredAction> constructor() { super('skypod') this.version(2).stores({ // ... existing stores ... actions: '×tamp', // Primary key is timestamp (unique due to LC) }) } } #+END_SRC *** 5. Action Dispatcher (`src/client/actions/dispatcher.ts`) #+BEGIN_SRC typescript import {SkypodDatabase} from '#client/database' import {RealmConnection} from '#client/realm/connection' import {LogicalClock, LCTimestamp, SkypodAction, skypodActionSchema} from '#common/protocol' import {ActionReducer} from './reducer' export class ActionDispatcher { private reducer: ActionReducer private realmConnection?: RealmConnection constructor( private db: SkypodDatabase, private logicalClock: LogicalClock ) { this.reducer = new ActionReducer(db) } setRealmConnection(connection: RealmConnection | undefined) { this.realmConnection = connection } async dispatch(action: SkypodAction): Promise<void> { // Validate action shape const parsed = skypodActionSchema.parse(action) // Add timestamp if not present (remote actions already have one) const actionWithTimestamp = parsed.timestamp ? parsed : { ...parsed, timestamp: await this.logicalClock.now() } // Apply locally through reducer await this.reducer.reduce(actionWithTimestamp) // Broadcast to peers if connected and not a remote action if (this.realmConnection?.connected && !action.meta?.skipSync) { this.realmConnection.broadcast({ typ: 'evt', msg: 'skypod.sync', dat: actionWithTimestamp }) } } // Get latest action timestamp for sync async getLatestTimestamp(): Promise<LCTimestamp | undefined> { const latest = await this.db.actions .orderBy('timestamp') .reverse() .first() return latest?.timestamp } // Get actions since a timestamp (for sync) async getActionsSince(since?: LCTimestamp): Promise<StoredAction[]> { if (!since) { return await this.db.actions .orderBy('timestamp') .toArray() } return await this.db.actions .where('timestamp') .above(since) .toArray() } // Apply batch of remote actions async syncActions(actions: SkypodAction[]): Promise<void> { // Sort by timestamp to apply in correct order const sorted = actions.sort((a, b) => LogicalClock.compare(a.timestamp!, b.timestamp!) ) for (const action of sorted) { await this.dispatch({ ...action, meta: { ...action.meta, skipSync: true } }) } } } #+END_SRC *** 6. Reducer Implementation (`src/client/actions/reducer.ts`) #+BEGIN_SRC typescript import {nanoid} from 'nanoid' import {SkypodDatabase} from '#client/database' import {LogicalClock, SkypodAction} from '#common/protocol' export class ActionReducer { constructor(private db: SkypodDatabase) {} async reduce(action: SkypodAction): Promise<void> { const timestamp = action.timestamp! // Store the action in history await this.db.actions.put({ ...action, timestamp, }) // Apply the action to state switch (action.type) { case 'channel:subscribe': { const {url, title} = action.payload // Check if already exists const existing = await this.db.feeds.where('url').equals(url).first() if (existing) { // Use LogicalClock comparison for conflict resolution if (existing.timestamp && LogicalClock.compare(timestamp, existing.timestamp) <= 0) { return // Existing is newer or same, skip } } // Add or update channel await this.db.feeds.put({ url, guid: existing?.guid || nanoid(), title: title || existing?.title, tags: [], refreshHP: 100, timestamp, // Store for future conflict resolution }) break } case 'channel:update': { const {guid, ...updates} = action.payload const existing = await this.db.feeds.where('guid').equals(guid).first() if (!existing) return // Can't update non-existent channel // Check timestamp for conflict resolution if (existing.timestamp && LogicalClock.compare(timestamp, existing.timestamp) <= 0) { return // Existing is newer } await this.db.feeds.where('guid').equals(guid).modify({ ...updates, timestamp, }) break } case 'channel:unsubscribe': { const {guid} = action.payload await this.db.feeds.where('guid').equals(guid).delete() break } case 'entry:markRead': { // Similar pattern for entries break } } } } #+END_SRC *** 7. Database Context Integration (`src/client/context-database.tsx`) #+BEGIN_SRC typescript import {ActionDispatcher} from './actions/dispatcher' import {RealmIdentityContext} from './realm/context-identity' import {RealmConnectionContext} from './realm/context-connection' import {SkypodAction} from '#common/protocol/skypod-actions' export interface SkypodDbContextValue { useDbSignal: <T>(querier: DbQuerier<T>) => ReadonlySignal<T | undefined> dispatch: (action: SkypodAction) => Promise<void> } export const SkypodDbProvider: preact.FunctionComponent<{children: preact.ComponentChildren}> = ( props, ) => { const db = useMemo(() => new SkypodDatabase(), []) const identityContext = useContext(RealmIdentityContext) const connectionContext = useContext(RealmConnectionContext) // Create dispatcher with logical clock from identity const dispatcher = useMemo(() => { if (!identityContext) return null return new ActionDispatcher(db, identityContext.clock) }, [db, identityContext]) // Update dispatcher's realm connection when it changes useEffect(() => { dispatcher?.setRealmConnection(connectionContext?.realm.value) }, [dispatcher, connectionContext?.realm.value]) // Listen for incoming sync messages useEffect(() => { const connection = connectionContext?.realm.value if (!connection || !dispatcher) return const handleMessage = async (event: CustomEvent) => { const message = event.detail if (message.msg === 'skypod.sync') { switch (message.typ) { case 'evt': // Real-time action broadcast await dispatcher.dispatch({ ...message.dat, meta: { ...message.dat.meta, skipSync: true } }) break case 'req': // Peer requesting sync const actions = await dispatcher.getActionsSince(message.dat.since) connection.send(message.source, { typ: 'res', msg: 'skypod.sync', seq: message.seq, dat: { actions } }) break case 'res': // Received sync response await dispatcher.syncActions(message.dat.actions) break } } } connection.addEventListener('message', handleMessage) return () => connection.removeEventListener('message', handleMessage) }, [connectionContext?.realm.value, dispatcher]) // Request sync when connecting to peers useEffect(() => { const connection = connectionContext?.realm.value if (!connection || !dispatcher) return const handlePeerConnected = async (event: CustomEvent) => { const latestTimestamp = await dispatcher.getLatestTimestamp() // Request sync from peer connection.send(event.detail.peerId, { typ: 'req', msg: 'skypod.sync', dat: { since: latestTimestamp } }) } connection.addEventListener('peer:connected', handlePeerConnected) return () => connection.removeEventListener('peer:connected', handlePeerConnected) }, [connectionContext?.realm.value, dispatcher]) function useDbSignal<T>(querier: DbQuerier<T>): ReadonlySignal<T | undefined> { // ... existing implementation } return ( <SkypodDbContext.Provider value={{ useDbSignal, dispatch: dispatcher ? dispatcher.dispatch.bind(dispatcher) : async () => {} }}> {props.children} </SkypodDbContext.Provider> ) } #+END_SRC *** 8. Component Usage Example #+BEGIN_SRC typescript // In webrtc-demo.tsx const {useDbSignal, dispatch} = useSkypodDb() const feeds = useDbSignal((db) => db.feeds.toArray()) const doSubscribe = useCallback(async () => { await dispatch({ type: 'channel:subscribe', payload: { url: subscribeUrl.value, } }) }, [subscribeUrl, dispatch]) #+END_SRC ** Key Design Decisions *** 1. LogicalClock Integration - Created per identity in RealmIdentityContext - Persists state to localStorage to survive reloads - Generates unique timestamps for total ordering *** 2. Action Flow - All state changes go through: dispatch → reducer → Dexie - Actions are validated with zod schemas - Timestamps added at dispatch time - Stored in actions table for sync *** 3. Sync Protocol - Uses existing realm WebRTC connections - Single message type `skypod.sync` for all operations - Request/response for initial sync - Events for real-time updates - Incremental sync using "since" timestamp *** 4. Conflict Resolution - LogicalClock.compare() determines winner - Last-write-wins semantics - Store timestamp with entities for future comparisons - Idempotent reducers handle out-of-order delivery *** 5. Database Context - Encapsulates db instance - Provides dispatch method - Handles realm connection lifecycle - Manages sync protocol messages ** Implementation Order 1. Create action schemas in protocol (actions.ts, skypod-actions.ts) 2. Add sync messages to protocol/messages.ts 3. Update database schema with actions table and timestamps 4. Implement ActionReducer with channel:subscribe 5. Implement ActionDispatcher with LogicalClock 6. Update SkypodDbProvider to create dispatcher and handle sync 7. Update WebRTC demo to use dispatch instead of direct db writes 8. Test P2P sync between browser tabs ** Future Enhancements *** Optimization - Batch actions for better performance - Compress action payloads - Implement action compaction for old events *** Features - Add more action types (entries, playback state, etc) - Implement undo/redo using action history - Add action replay for debugging - Create Redux DevTools integration *** Scaling - Implement checkpoints for faster initial sync - Add pagination for large action histories - Consider moving old actions to OPFS - Implement selective sync (by feed, date range, etc)