diff --git a/content/blog/2025-05-01.md b/content/blog/2025-05-01.md index b42b227..28c5e78 100644 --- a/content/blog/2025-05-01.md +++ b/content/blog/2025-05-01.md @@ -1,11 +1,11 @@ --- -title: "Easy cloud compute for Comind" +title: "Easy cloud compute for Comind (Coming Soon)" date: 2025-05-01T16:43:23-07:00 description: "I started working on some convenience tools for using cloud compute to power a Comind instance. It's not quite done, but should be useful down the line." draft: false --- -**Claude summary:** This post introduces a new way to run Comind in the cloud using Modal and vLLM. We've created simple deployment scripts that handle API key management, container warming, and server deployment. If you don't have access to powerful GPUs locally, you can now easily deploy Comind to Modal's cloud infrastructure with just a few commands. +**Claude summary:** This post introduces a new way to run Comind in the cloud using Modal and vLLM that will soon be available. We're creating simple deployment scripts that will handle API key management, container warming, and server deployment. If you don't have access to powerful GPUs locally, you will soon be able to easily deploy Comind to Modal's cloud infrastructure with just a few commands. ## Cameron's note @@ -18,15 +18,15 @@ I vibe code Comind quite a bit. I'm going to have Claude write up our changes in A big issue with Comind is that it requires a set of structured output features that are not supported by commercial providers, so you have to run the model yourself. Most people don't have a giant GPU like I do, so I wanted to provide a simple way to run the model in the cloud. -I chose to use [Modal](https://modal.com), a straightforward and easy service to deploy a [vLLM](https://github.com/vllm-project/vllm) server. vLLM is probably the best inference server available, and has a tight integration with several structured output libraries. An additional benefit of using Modal is that it's easy to deploy a server that can be accessed by any client that supports the [OpenAI API](https://platform.openai.com/docs/api-reference/introduction). +I've chosen to use [Modal](https://modal.com), a straightforward and easy service to deploy a [vLLM](https://github.com/vllm-project/vllm) server. vLLM is probably the best inference server available, and has a tight integration with several structured output libraries. An additional benefit of using Modal is that it will be easy to deploy a server that can be accessed by any client that supports the [OpenAI API](https://platform.openai.com/docs/api-reference/introduction). -## Updated Modal Deployment Process +## Upcoming Modal Deployment Process -After addressing some compatibility issues with Modal, here's the updated deployment process: +I'm currently addressing some compatibility issues with Modal. Here's the planned deployment process that will be available soon: ### Option 1: Quick Interactive Setup (Recommended) -Use our deployment helper script that will guide you through the process: +You'll be able to use our deployment helper script that will guide you through the process: ```bash # Deploy and setup interactively (recommended) @@ -41,7 +41,7 @@ python modal_deploy.py status ### Option 2: Manual Setup -If you prefer to set things up manually: +If you prefer to set things up manually, you'll be able to: 1. **Create a Secret in Modal** (optional but recommended): ```bash @@ -65,15 +65,15 @@ If you prefer to set things up manually: ### Troubleshooting -If you encounter deployment errors: +If you encounter deployment errors once this feature is released: -1. **Secret not found**: You can either: +1. **Secret not found**: You'll be able to either: - Create the secret as shown above - Continue without a secret (a default API key will be used) -2. **Deprecation warnings**: These are informational for future Modal updates and won't affect functionality currently. +2. **Deprecation warnings**: These will be informational for future Modal updates and shouldn't affect functionality. -3. **Authorization errors**: Make sure your client is using the same API key as your server: +3. **Authorization errors**: You'll need to make sure your client is using the same API key as your server: ```python client = OpenAI( api_key="your-api-key-here", # Must match what you set in Modal @@ -81,7 +81,7 @@ If you encounter deployment errors: ) ``` -Now, copy these URLs into the `.env` file for your Comind instance: +Once deployed, you'll be able to copy these URLs into the `.env` file for your Comind instance: ```bash # LLM server info @@ -94,7 +94,7 @@ COMIND_EMBEDDING_SERVER_API_KEY= "comind-api-key" ``` > [!NOTE] -> The inference server uses the API key `comind-api-key` by default. You can change this in the `modal_client.py` script: +> The inference server will use the API key `comind-api-key` by default. You'll be able to change this in the `modal_client.py` script: > > ```python > # Configuration options (can be modified as needed) @@ -103,7 +103,7 @@ COMIND_EMBEDDING_SERVER_API_KEY= "comind-api-key" > API_KEY = "comind-api-key" # Replace with a secret for production use > ``` -After deployment, you can access your models through the OpenAI client: +After deployment, you'll be able to access your models through the OpenAI client: ```python from openai import OpenAI @@ -124,22 +124,22 @@ response = client.chat.completions.create( (it will not know what Comind is for sure) -I've also created a simple client script that you can use to test the inference server: +I'm also working on a simple client script that you'll be able to use to test the inference server: ```bash python modal_client.py --workspace YOUR_WORKSPACE --prompt "Tell me about Comind" ``` -It currently only supports Phi-4. PRs welcome to add more models! I did a weird job so please help. +The initial version will only support Phi-4. PRs will be welcome to add more models! > [!NOTE] -> Keep in mind that the server has a warmup time, so it may take a while for it to boot up. +> Keep in mind that the server will have a warmup time, so it may take a while for it to boot up. -## Solving Cold Start Problems +## Addressing Cold Start Problems -One common issue with Modal and similar serverless platforms is cold start time - the delay when a new container needs to be initialized. For a quick fix: +One common issue with Modal and similar serverless platforms is cold start time - the delay when a new container needs to be initialized. The planned solutions include: -1. **Keep containers warm** by adding these parameters to your Modal functions: +1. **Keeping containers warm** by adding these parameters to Modal functions: ```python @app.function( @@ -152,7 +152,7 @@ One common issue with Modal and similar serverless platforms is cold start time ) ``` -2. **Update immediately after deployment** with: +2. **Updating immediately after deployment** with: ```python if __name__ == "__main__": @@ -162,7 +162,7 @@ if __name__ == "__main__": serve_phi4.keep_warm(1) ``` -3. **Schedule warm container adjustments** based on time of day: +3. **Scheduling warm container adjustments** based on time of day: ```python @app.function(schedule=modal.Cron("0 * * * *")) @@ -174,9 +174,9 @@ def adjust_warm_containers(): serve_phi4.keep_warm(1) ``` -## Authorization Errors +## Preventing Authorization Errors -If you're seeing authorization errors, make sure: +To prevent authorization errors, you'll need to make sure: 1. The API key in your client matches the server: ```python @@ -196,7 +196,7 @@ If you're seeing authorization errors, make sure: ## Easier Deployment with modal_deploy.py -I've also created a deployment helper script that simplifies the process and solves the cold start issues automatically: +I'm also creating a deployment helper script that will simplify the process and solve cold start issues automatically: ```bash # Deploy and keep containers warm in one step @@ -209,22 +209,22 @@ python modal_deploy.py warm python modal_deploy.py status ``` -This script automatically keeps containers warm after deployment and shows you your endpoint URLs based on your Modal workspace name. It's a much better experience than the manual deployment process. +This script will automatically keep containers warm after deployment and show you your endpoint URLs based on your Modal workspace name, providing a much better experience than the manual deployment process. ## Securing API Keys -Instead of hardcoding API keys (which is never a good idea), the updated version uses Modal's built-in secret management: +Instead of hardcoding API keys (which is never a good idea), the upcoming version will use Modal's built-in secret management: ```bash # Create a secure API key (run this once) modal secret create comind-api-key --value "your-secure-key-here" ``` -The deployment script automatically checks if this secret exists and creates it with a default value if needed. This provides three benefits: +The deployment script will automatically check if this secret exists and create it with a default value if needed. This will provide three benefits: -1. Your API key isn't stored in source code -2. You can rotate keys without changing code -3. The same key is consistently used across all services +1. Your API key won't be stored in source code +2. You'll be able to rotate keys without changing code +3. The same key will be consistently used across all services In your client code, you'll use this same key: @@ -235,16 +235,16 @@ client = OpenAI( ) ``` -This eliminates the "unauthorized" errors that happen when keys don't match between client and server. +This will eliminate the "unauthorized" errors that happen when keys don't match between client and server. ## Handling Modal Secrets Properly > [!IMPORTANT] -> There's an important update regarding Modal secrets handling. If you're seeing errors like `AttributeError: 'Secret' object has no attribute 'get'` or `TypeError: _App.function() got an unexpected keyword argument 'env'`, the code has been updated to fix these issues. +> There's an important note regarding Modal secrets handling that will be implemented. If you see errors like `AttributeError: 'Secret' object has no attribute 'get'` or `TypeError: _App.function() got an unexpected keyword argument 'env'`, the code will include solutions to these issues. -Modal's Secret API works differently than we initially expected. Here's the correct way to use Modal secrets: +Modal's Secret API works differently than initially expected. Here's how Modal secrets will be implemented: -1. **Create a secret**: +1. **Creating a secret**: ```python # Create a secret from a dictionary api_key_secret = modal.Secret.from_dict({"api_key": "comind-api-key"}) @@ -253,7 +253,7 @@ Modal's Secret API works differently than we initially expected. Here's the corr api_key_secret = modal.Secret.from_name("comind-api-key") ``` -2. **Pass the secret to your functions**: +2. **Passing the secret to functions**: ```python @app.function( image=vllm_image, @@ -264,7 +264,7 @@ Modal's Secret API works differently than we initially expected. Here's the corr # function code... ``` -3. **Access the secret in your function**: +3. **Accessing the secret in functions**: ```python def get_api_key(): """Get the API key from the environment.""" @@ -273,19 +273,19 @@ Modal's Secret API works differently than we initially expected. Here's the corr return os.environ.get("api_key", "comind-api-key") ``` -The secret values are injected as environment variables in your container, so you access them with `os.environ`. This pattern is now implemented in all the Modal functions in our codebase. +The secret values will be injected as environment variables in your container, so you'll access them with `os.environ`. This pattern will be implemented in all the Modal functions in our codebase. -## Current Status and Known Issues +## Current Development Status -As of the latest update, there are still some ongoing issues with the Modal interface that we're actively working to resolve: +As we prepare to release this feature, we're working on resolving several issues with the Modal interface: -1. **API Integration Issues**: Some users are experiencing inconsistent responses when connecting their Comind instance to the Modal-hosted inference server. We're investigating the root cause, which appears to be related to how the API endpoints handle certain request formats. +1. **API Integration Issues**: We're developing solutions for consistent responses when connecting Comind instances to Modal-hosted inference servers, focusing on how API endpoints handle various request formats. -2. **Container Warmup Reliability**: Despite the warmup mechanisms we've implemented, some users may still experience occasional cold start delays. We're fine-tuning the container management logic to improve reliability. +2. **Container Warmup Reliability**: We're fine-tuning container management logic to minimize cold start delays, ensuring a more responsive experience. -3. **Authentication Edge Cases**: In certain scenarios, authentication between the client and server may fail even with correctly configured API keys. We're working on more robust error handling to make these cases more diagnosable. +3. **Authentication Edge Cases**: We're implementing more robust error handling for authentication between clients and servers to make troubleshooting easier. -If you encounter any of these issues, please help us improve by reporting specific error messages and the steps to reproduce in the project's issue tracker. We're actively monitoring and addressing these concerns to make the cloud deployment experience as seamless as possible. +Once these issues are resolved, we'll release the Modal cloud deployment feature. We look forward to your feedback and contributions when this functionality becomes available. -- Cameron