Local Llm Ops

by bobmatnyc718070a7d622MIT77 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 2 months ago

Local LLM operations with Ollama on Apple Silicon, including setup, model pulls, chat launchers, benchmarks, and diagnostics.

Instructions onlyDevOps & Cloud
AI-generated overview

Operational guide for running local LLMs with Ollama on Apple Silicon, covering setup, model pulls, chat, benchmarks and diagnostics.

What it does
This skill documents an operational workflow for running local large language models with Ollama on Apple Silicon. It walks through installing Ollama, starting the service, initializing a virtual environment, pulling models, and launching chat sessions or benchmark runs. It also covers diagnostic steps and common fixes when setup fails.
When to use it
Use it when operating local LLMs on macOS, running Ollama-based chat sessions, or benchmarking models for speed and latency. It is also relevant when troubleshooting a local Ollama setup that is not working.
Requirements
Requires Ollama installed and running as a service on macOS, a shell environment with the referenced setup, chat, benchmark and diagnostic scripts, and network access to pull models. The skill itself is instructions only and ships no scripts.

Local LLM Ops (Ollama)

Overview

Your localLLM repo provides a full local LLM toolchain on Apple Silicon: setup scripts, a rich CLI chat launcher, benchmarks, and diagnostics. The operational path is: install Ollama, ensure the service is running, initialize the venv, pull models, then launch chat or benchmarks.

Quick Start

bash
./setup_chatbot.sh./chatllm

If no models are present:

bash
ollama pull mistral

Setup Checklist

  1. Install Ollama: brew install ollama
  2. Start the service: brew services start ollama
  3. Run setup: ./setup_chatbot.sh
  4. Verify service: curl http://localhost:11434/api/version

Chat Launchers

  • ./chatllm (primary launcher)
  • ./chat or ./chat.py (alternate launchers)
  • Aliases: ./install_aliases.sh then llm, llm-code, llm-fast

Task modes:

bash
./chat -t coding -m codellama:70b./chat -t creative -m llama3.1:70b./chat -t analytical

Benchmark Workflow

Benchmarks are scripted in scripts/run_benchmarks.sh:

bash
./scripts/run_benchmarks.sh

This runs bench_ollama.py with:

  • benchmarks/prompts.yaml
  • benchmarks/models.yaml
  • Multiple runs and max token limits

Diagnostics

Run the built-in diagnostic script when setup fails:

bash
./diagnose.sh

Common fixes:

  • Re-run ./setup_chatbot.sh
  • Ensure ollama is in PATH
  • Pull at least one model: ollama pull mistral

Operational Notes

  • Virtualenv lives in .venv
  • Chat configs and sessions live under ~/.localllm/
  • Ollama API runs at http://localhost:11434

Related Skills

  • toolchains/universal/infrastructure/docker

Source and attribution

Source:bobmatnyc/claude-mpm-skillsintoolchains/ai/ops/local-llm-opsat commit718070a

License: MIT

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal