# Branch: py4lexis4 ## Epic Reference **Epic**: Py4Lexis 5.x Migration **Goal**: Migrate owilix to py4lexis 5.x with HTTP iRODS client --- ## Problem Statement Py4lexis 5.x replaces `python-irodsclient` with `irods_http_client`. This breaks: - Current fsspec integration (`irods-fsspec` package) - Session/token management in OWILIXManager - File access patterns in repositories --- ## Cycles Overview | Cycle | Goal | Status | |-------|------|--------| | 1 | Create iRODS HTTP fsspec client | βœ… Complete | | 1a | Async optimization (httpx) | βœ… Complete | | 2 | Session management & config refactor | 🟑 Planning | | 3 | Repository integration (HTTP2IRODSLexisRepository) | βšͺ Planned | | 4 | Code cleanup analysis | βšͺ Planned | --- ## Cycle 1: iRODS HTTP fsspec Client **Goal**: Create `owilix.core.fsspec` with sync/async fsspec wrapper for `irods_http_client`. ### Key Design Decisions - **Protocol name**: `http2irods` - **Async support**: Full `AsyncFileSystem` capability for future async owilix - **Performance focus**: HTTP/2 multiplexing, connection pooling, configurable chunk sizes ### Implementation #### [NEW] `owilix/core/fsspec/__init__.py` Package exposing `Http2IrodsFileSystem`. #### [NEW] `owilix/core/fsspec/http2irods.py` ```python class Http2IrodsFileSystem(fsspec.AbstractFileSystem): """fsspec filesystem for iRODS HTTP API.""" protocol = "http2irods" async_impl = True # Enable async capabilities def __init__(self, session, url_base, chunk_size=1024*1024, # Configurable chunk size max_connections=10, # Connection pool size **kwargs): ... # Sync methods def _open(self, path, mode="rb", **kwargs) -> Http2IrodsFile def ls(self, path, detail=False) -> list def info(self, path) -> dict # Async methods (for future async owilix) async def _cat_file(self, path, start=None, end=None) async def _get_file(self, rpath, lpath, **kwargs) async def _ls(self, path, detail=False) ``` #### [NEW] `owilix/core/fsspec/http2irods_file.py` ```python class Http2IrodsFile(fsspec.spec.AbstractBufferedFile): """File handle with range request support.""" def _fetch_range(self, start, end): """Use DataObjects.read(path, offset, count).""" ``` ### Performance Considerations - [ ] **HTTP/2 support**: Check if `irods_http_client` uses httpx/aiohttp with HTTP/2 - [ ] **Connection pooling**: Keep connections open for multiple requests - [ ] **Chunk size tuning**: Benchmark different sizes (64KB, 256KB, 1MB, 4MB) - [ ] **Offset alignment**: Test if aligned offsets improve performance ### Benchmarks ``` tests/owilix/core/fsspec/ β”œβ”€β”€ benchmark_core_fsspec_chunk_sizes.py # Compare chunk sizes β”œβ”€β”€ benchmark_core_fsspec_connections.py # Connection pool impact └── benchmark_core_fsspec_sequential.py # Sequential vs parallel reads ``` ### Tests ``` tests/owilix/core/fsspec/ β”œβ”€β”€ test_core_fsspec_unit.py # Mocked unit tests β”œβ”€β”€ test_core_fsspec_integration.py # @pytest.mark.integration └── example_core_fsspec_usage.py # Usage examples ``` ### Verification - [ ] Unit tests pass - [ ] Integration tests with dataset `f6ea5756-2e0b-11ef-b336-0242ac1d0004` - [ ] Benchmark results documented --- ## Cycle 2: Session Management & Config Refactor βœ… **Goal**: Refactor `OWILIXManager` and `OWILIXConfig` into a clean package. Ensure robust session handling. **Status**: Complete ### Implemented Changes 1. **Package Restructure**: Created `owilix/core/manager/` package. 2. **Pydantic Configuration**: `OWILIXSettings`, `RepositoryConfig`, etc. 3. **Session Wrapper**: `OWILIXSession` with lazy-loaded iRODS. 4. **UI Integration**: Moved `ui.py` into manager package. 5. **Optional Dependencies**: Added `irods-tcp` extras group. 6. **Tests**: 12 unit tests for env, config, session, backward compat. ### Final Package Structure ``` owilix/core/manager/ β”œβ”€β”€ __init__.py # Package exports β”œβ”€β”€ env.py # OWILIXEnv singleton β”œβ”€β”€ config.py # Pydantic config models β”œβ”€β”€ session.py # OWILIXSession wrapper β”œβ”€β”€ manager.py # OWILIXManager class └── ui.py # Console and progress utilities ``` ### Backward Compatibility Old import paths still work via re-export shims: - `owilix.core.env` β†’ `owilix.core.manager.env` - `owilix.core.ui` β†’ `owilix.core.manager.ui` - `owilix.core.manager` (file) β†’ `owilix.core.manager` (package) --- ## Cycle 3: Repository Integration **Goal**: Create `HTTP2IRODSLexisRepository` - a new repository class. ### Architecture - **Derive from** `FileBasedRepository` (not `IRODSRepository`). - **Metadata**: Use `OWILexisDatasetAPI` (LEXIS DDI HTTP API). - **Files**: Use `Http2IrodsFileSystem` (iRODS HTTP API via fsspec). - **Keep** `LEXISIrodsHTTPRepository` as legacy (optional dependency). ### Key Changes 1. Clean separation: DDI API for metadata ↔ iRODS HTTP API for files. 2. No `python-irodsclient` dependency in new repository. 3. Refactor `repository.py` for cleaner inheritance. ### Files to Create/Modify | Action | File | Description | |--------|------|-------------| | NEW | `core/repository/http2irods.py` | `HTTP2IRODSLexisRepository` | | MODIFY | `core/repository.py` | Refactor base classes | --- ## Cycle 4: Code Cleanup Analysis **Goal**: Analyze owilix↔py4lexis integration, identify redundant code. ### Deliverables - Analysis document of current integration - List of redundant/removable code - Recommendations for improvements - Additional cycles if major refactoring needed --- ## Key Analysis Notes ### From small_tst_py4lexis.py ```python # Token refresh requires wrapping class OWIIrods(iRODS): def irods(self): self._iRODS__check_access_token() # Must call before each operation return self._irds # Range requests DataObjects(irods, url_base).read(path, offset=10, count=20) # Async support available coll.data_objects[0].get_async_reader() ``` --- ## irods_http_client API Reference ### Collection Operations (`collection_operations.py`) | Method | Parameters | Description | |--------|------------|-------------| | `list(lpath, recurse=0/1)` | Single HTTP call | **Fast**: recurse=1 returns all entries recursively | | `stat(lpath)` | | Collection metadata (permissions, timestamps) | | `remove(lpath, recurse, no_trash)` | | Delete collection | | `create(lpath, create_intermediates)` | | Create collection | ### iRODSCollection Model (`models/collection.py`) ```python class iRODSCollecion: # Properties (use GenQuery with pagination, 256 items/page) subcollections # List of subcollections data_objects # List of data objects metadata # Collection metadata # Methods walk(topdown=True) # Generator like os.walk(), yields (coll, subcols, objects) stat() # Calls Collection.stat() remove(recurse) # Calls Collection.remove() ``` ### Data Object Operations (`data_object_operations.py`) | Method | Parameters | Description | |--------|------------|-------------| | `read(lpath, offset, count)` | Range request | Read bytes at offset | | `stat(lpath)` | | File size, checksum, timestamps | | `get(lpath)` | | Returns iRODSDataObject | | `exists(lpath)` | | Check if file exists | | `parallel_read(lpath, local, workers, chunk_size)` | | Multithreaded download | | `parallel_write(local, lpath, workers, chunk_size)` | | Multithreaded upload | ### iRODSDataObject Model (`models/data_object.py`) ```python class iRODSDataObject: size # File size in bytes path # Full iRODS path checksum # File checksum open(mode='r') # Returns file-like iRODSDataObjectFileRaw get_async_reader(worker_count, chunk_size) # Parallel chunk reader get_async_writer(worker_count) # Parallel chunk writer ``` ### Performance Recommendations | Operation | Preferred Method | Reason | |-----------|------------------|--------| | **List files** | `Collections.list(recurse=1)` | Single HTTP call vs multiple | | **Large download** | `get_async_reader(workers=5)` | Parallel HTTP requests | | **Range reads** | `DataObjects.read(offset, count)` | Direct HTTP with Range header | | **Traversal** | `iRODSCollection.walk()` | Memory efficient iterator | ### Benchmark Results (Updated 2026-01-01) | Metric | Best Config | Result | |--------|-------------|--------| | **File Listing** | `Collections.list(recurse=1)` | **5.2s** for 1893 files | | **Sequential Read** | chunk_size=1MB | 3.18 MB/s | | **DuckDB count(*)** | 1MB chunk | **1.3s** | | **DuckDB Schema** | 1MB chunk | **1.1s** (DESCRIBE) | | **DuckDB Column** | 1MB chunk | **3.5s** (SELECT id LIMIT 5000) | #### Optimization Notes 1. **Listing**: Always use `Collections.list(recurse=1)`. It is fast and efficient. 2. **DuckDB**: Fully functional. `mtime` and protocol stripping implemented. 3. **Chunk Size**: 1MB is the sweet spot. 256KB is slightly slower for queries; 4MB offers no benefit. ### 6. Async & Concurrency Optimization (Cycle 1a) **Findings**: - **Protocol**: The LEXIS iRODS server currently supports **HTTP/1.1** (no HTTP/2 multiplexing). - **Performance**: - **Sequential Async**: Lower performance (~0.7 MB/s) due to lack of multiplexing and HTTP latency on small files. - **Concurrent Async**: **Significantly faster (~4.45 MB/s)** when fetching multiple files in parallel (using `asyncio.gather`), outperforming synchronous sequential reads (~3.2 MB/s). **Recommendation**: - Use `AsyncHttp2IrodsFileSystem` with `httpx` for scenarios involving **concurrent file access** (e.g., scraping many small files). - For single large file streaming, the synchronous implementation (via `rows` or `requests`) remains robust. - Enable `httpx` connection pooling (`limits=httpx.Limits(max_keepalive_connections=20, max_connections=50)`). For detailed usage, architecture, and performance breakdowns, see [FSSPEC Integration](../source/fsspec_integration.md).