GrepAI Embeddings with LM Studio
This skill covers using LM Studio as the embedding provider for GrepAI, offering a user-friendly GUI for managing local models.
When to Use This Skill
- Want local embeddings with a graphical interface
- Already using LM Studio for other AI tasks
- Prefer visual model management over CLI
- Need to easily switch between models
What is LM Studio?
LM Studio is a desktop application for running local LLMs with:
- 🖥️ Graphical user interface
- 📦 Easy model downloading
- 🔌 OpenAI-compatible API
- 🔒 100% private, local processing
Prerequisites
- Download LM Studio from lmstudio.ai
- Install and launch the application
- Download an embedding model
Installation
Step 1: Download LM Studio
Visit lmstudio.ai and download for your platform:
- macOS (Intel or Apple Silicon)
- Windows
- Linux
Step 2: Launch and Download a Model
- Open LM Studio
- Go to the Search tab
- Search for an embedding model:
nomic-embed-text-v1.5bge-small-en-v1.5bge-large-en-v1.5
- Click Download
Step 3: Start the Local Server
- Go to the Local Server tab
- Select your embedding model
- Click Start Server
- Note the endpoint (default:
http://localhost:1234)
Configuration
Basic Configuration
With Custom Port
With Explicit Dimensions
Available Models
nomic-embed-text-v1.5 (Recommended)
bge-small-en-v1.5
Best for: Smaller codebases, faster indexing.
bge-large-en-v1.5
Best for: Maximum accuracy.
Model Comparison
LM Studio Server Setup
Starting the Server
- Open LM Studio
- Navigate to Local Server tab (left sidebar)
- Select an embedding model from the dropdown
- Configure settings:
- Port:
1234(default) - Enable Embedding Endpoint
- Port:
- Click Start Server
Server Status
Look for the green indicator showing the server is running.
Verifying the Server
LM Studio Settings
Recommended Settings
In LM Studio's Local Server tab:
GPU Acceleration
LM Studio automatically uses:
- macOS: Metal (Apple Silicon)
- Windows/Linux: CUDA (NVIDIA)
Adjust GPU layers in settings for memory/speed balance.
Running LM Studio Headless
For server environments, LM Studio supports CLI mode:
Common Issues
❌ Problem: Connection refused ✅ Solution: Ensure LM Studio server is running:
- Open LM Studio
- Go to Local Server tab
- Click Start Server
❌ Problem: Model not found ✅ Solution:
- Download the model in LM Studio's Search tab
- Select it in the Local Server dropdown
❌ Problem: Slow embedding generation ✅ Solutions:
- Enable GPU acceleration in LM Studio settings
- Use a smaller model (bge-small-en-v1.5)
- Close other GPU-intensive applications
❌ Problem: Port already in use ✅ Solution: Change port in LM Studio settings:
❌ Problem: LM Studio closes and server stops ✅ Solution: Keep LM Studio running in the background, or consider using Ollama which runs as a system service
LM Studio vs Ollama
Recommendation: Use LM Studio if you prefer a GUI, Ollama for always-on background service.
Migrating from LM Studio to Ollama
If you need a more reliable background service:
- Install Ollama:
- Update config:
- Re-index:
Best Practices
- Keep LM Studio running: Server stops when app closes
- Use recommended model:
nomic-embed-text-v1.5for best balance - Enable GPU: Faster embeddings with hardware acceleration
- Check server before indexing: Ensure green status indicator
- Consider Ollama for production: More reliable as background service
Output Format
Successful LM Studio configuration:


