# Hermes Ollama Models Native-like Hermes dashboard plugin for local Ollama model management and chat. ## Chat capabilities - Select an installed Ollama model and load it into memory - Chat through Ollama's native `/api/chat` endpoint - Attach screenshots and JPEG/PNG/WebP images for vision-capable models - Attach text PDFs; PDF text is extracted with `pypdf` - Add public HTTP/HTTPS URLs for HTML/text, images, or PDFs - View live host RAM and swap statistics - View Ollama's loaded-model memory split: total, GPU VRAM, and normal RAM/offload - View NVIDIA GPU telemetry when `nvidia-smi` is available - Chat is the default view when the Ollama Models plugin opens - Natural composer behavior: Enter sends; Shift+Enter creates a new line - Paste images directly into the composer and drag/drop images, PDFs, and text files - Streamed Ollama responses with a real Stop action that cancels the active request - Minimized-by-default expandable thinking/progress details with live stage, elapsed time, event, and character counters The chat transcript and selected model are persisted in the Hermes server's SQLite database at `~/.hermes/ollama-manager/chat.sqlite3`, so conversations can be listed and resumed from another browser or after a dashboard restart. Use **New conversation** to start a separate thread and **Clear chat** to delete the selected shared conversation. Uploaded files remain temporary; attachment names/types/URLs are retained as metadata, not raw file contents. ## Performance metrics Each completed request records Ollama-provided counters and timing when available: - Time to first token (TTFT) - Total request latency - Prompt/input token count and prompt tokens/sec - Output token count and output tokens/sec - Total tokens - Ollama total/load/prompt-evaluation/generation durations - Model, request status, request ID, and error details The dashboard shows per-conversation metrics and rolling aggregates across the latest 100 requests. Values are shown as unavailable when Ollama does not provide them; the plugin does not estimate token counts. ## Catalog and download filters The Popular and Available downloads views only show models with known size and expected RAM estimates that fit the running host's installed RAM, detected from `/proc/meminfo` and reported in the API filter metadata. The limit is calculated at runtime, so it follows RAM upgrades or downgrades without a hard-coded threshold. Available downloads also includes known fitting variants from installed model families. The Available downloads controls support: - Dense or MoE model type - Abilities such as completion, thinking, tools, vision, audio, and video when supplied by the catalog - Organization by upstream popularity, newest (`modified_at`), smallest size, largest size, or name - Search across model names, families, strengths, and capabilities Popularity and date ordering use upstream metadata only; the plugin does not invent popularity, dates, RAM requirements, or token metrics. ## Available catalog source Available downloads now combine Ollama's popular API response with the public `https://ollama.com/library` index and each public family tag page. This means models that are not currently popular and are not installed locally—such as `ornith-1.5:9b` and `ornith-1.5:35b`—are discoverable. The catalog refresh found 235 public family slugs and 7,230 raw variants during validation. The existing MLX exclusion and host-RAM fit filter still apply, so very large variants such as `ornith-1.5:397b` remain hidden when they cannot fit the detected host RAM. The Live runtime panel now shows overall CPU usage, logical CPU count, load averages, overall GPU utilization, and per-GPU VRAM usage. When multiple logical CPUs are detected, it expands into a scrollable responsive per-core grid. When multiple GPUs are detected, it expands into a responsive per-GPU grid showing utilization, VRAM used/free, temperature, and power when the driver reports them. The grids use auto-fit sizing and bounded scrolling so the panel scales to larger CPU and GPU counts without overflowing the dashboard. A single **Load selected permanently** action now loads models in verified stages. RAM-only models are attempted first, followed by GPU + RAM models. The backend checks Ollama `/api/ps` after each runner starts and automatically performs one recovery pass for any selected model Ollama evicted. Already resident models are not reloaded. The result includes a retry count and reports when automatic eviction recovery completed, so users do not need to click the load action again manually. Each installed model in the Model pool has a persistent placement selector: - **GPU + RAM (automatic offload)**: Ollama uses GPU layers where it can and keeps the remainder in system RAM. This is the default. - **RAM only (CPU)**: the plugin sends Ollama `num_gpu: 0`, preventing GPU layer offload for that model. Placement is sent both when loading models and when chatting, so a RAM-only model is not silently reloaded with GPU offload. RAM-only models will not increase GPU VRAM usage; GPU+RAM models can still be evicted by Ollama if the GPU/device-memory scheduler cannot fit the runner. When Ollama accepts a load request but evicts one model while starting another, the dashboard displays an amber capacity warning rather than a red plugin error. On the verified host, Ollama logged that a 27.9 GiB runner would exceed available device memory with approximately 2.5 GiB GPU memory free. The T600 has 4 GiB VRAM, so two large multimodal models cannot be guaranteed resident simultaneously by the plugin. The actual resident set remains authoritative through `/api/ps`. ## Live model-loading telemetry While Ollama is starting a runner, the highlighted runtime chart displays an animated **Loading into Ollama memory** state with the selected model names, current stage, and elapsed time. The runtime panel separately reports Ollama resident model-weight bytes and the estimated target weight bytes. This is separate from host `MemAvailable`: CPU-mapped model files may appear as Linux file cache rather than ordinary process RAM usage. ## Multi-model loading and resident state The model pool now verifies every load request against Ollama `/api/ps` before reporting success. The UI shows an in-progress loading message, then reports which models are actually resident and which Ollama evicted. Resident models are highlighted in the pool with a green loaded state. After a browser refresh, resident models repopulate the pool selection and the **Models for this answer** selector, allowing multiple loaded models to be selected for parallel chat. Ollama still controls the physical resident-model limit. If it cannot keep all requested models at once because of its scheduler, GPU policy, context allocation, or available memory, the plugin reports the non-resident names instead of claiming they were permanently loaded. Increasing that limit requires changing the Ollama service configuration; the plugin does not silently alter or restart the Ollama service. ## Hermes native model selector During plugin load after install/update, the plugin registers the active Ollama endpoint(s) in Hermes' native `providers:` configuration as **Ollama Models (Local)** and, when configured, **Ollama Models (Remote)**. Hermes discovers the installed IDs through each endpoint's OpenAI-compatible `/v1/models` route, so models such as `gemma4:latest` become selectable with their exact Ollama tags. Changing the Hermes model selector only changes the selected provider/model. It does not call Ollama load or keep-alive APIs. End users must use the Ollama Models page's **Load selected permanently** action when they want a model loaded and kept available. Installing/updating the plugin and synchronizing providers does not change Hermes' existing default model or provider. ## Installation This plugin is installable from the Hermes dashboard Plugin Section using the repository URL: ```text https://gitea.beyondcloud.solutions/dennii/Hermes-Ollama_Models.git ``` In Hermes Dashboard, open **Plugins**, choose **Install from repository**, enter the URL above, and install. The repository contains the root `plugin.yaml`, dashboard manifest, backend API, compiled frontend bundle, stylesheet, and an opt-in prerequisite declaration. On Linux, the Hermes installer will verify Ollama, install it with the official Ollama installer when missing, and install the plugin's `pypdf` dependency before committing the plugin into `~/.hermes/plugins/`. Ollama installation requires the Hermes container to run as root, which is the expected configuration for a privileged ZimaOS deployment. If the container is not running as root, the plugin install stops without enabling a partially configured plugin. After installation or an update, restart only the Hermes dashboard service if requested by the installer. The dashboard now includes a connection panel in the header. Enter an Ollama base URL, choose **Local** or **Remote**, and use **Test** before **Save**. Saved URLs are stored in Hermes plugin data with file permissions restricted to the Hermes account. Credential-bearing URLs are rejected. The plugin probes the configured local endpoint and known Docker endpoints, and tests a configured remote endpoint. A successful connection reports the Ollama version and installed model count. The active local endpoint continues to drive chat, status, load, and runtime telemetry. A URL saved from the dashboard takes precedence over the initial `OLLAMA_HOST` environment fallback, so the field can correct or replace an inherited container setting. When both distinct local and remote endpoints are online, downloading or re-downloading a model opens a destination chooser. The selected target is recorded on the job and the pull is sent to that endpoint. Remote downloads do not alter the local installed-model inventory. The plugin writes this configuration atomically and fixes the dashboard status-poll race that could let a slower older `/status` response overwrite a newer saved state. The config root is captured once per dashboard process from `HERMES_HOME`/Hermes home, rather than changing between requests. Removing the saved Local endpoint clears the active override and restores `OLLAMA_HOST` or automatic discovery. Removing Remote only clears the saved remote URL; detected-but-unsaved rows cannot be removed. If Local and Remote contain the same URL, the dashboard marks the duplicate and treats it as one download target; enter the actual second Ollama server URL to enable independent remote downloads. The plugin supports both deployment types: - **Physical/Linux Hermes:** Ollama defaults to `http://localhost:11434`. If Ollama is missing, the prerequisite installer can install it on a root-run host. - **Docker Hermes:** Ollama should normally run as a separate service or on the Docker host. The plugin does not install Ollama inside the Hermes container. It first probes `http://ollama:11434` and then `http://host.docker.internal:11434`. Set `OLLAMA_HOST` to the reachable Ollama base URL, without `/v1`, for example `http://ollama:11434`, `http://host.docker.internal:11434`, or the host address shown by your installation, such as `http://:11434`. For host access on Linux Docker, add this to the Hermes service when needed: ```yaml extra_hosts: - "host.docker.internal:host-gateway" environment: OLLAMA_HOST: http://host.docker.internal:11434 ``` For Compose service-to-service networking, use the Ollama service name instead: ```yaml environment: OLLAMA_HOST: http://ollama:11434 ``` The dashboard plugin API is mounted when the dashboard starts. Restart Hermes after installing or enabling the plugin; otherwise the plugin page can load while `/api/plugins/ollama-manager/status` still returns 404. ## Download storage The plugin does not store model blobs in Hermes. It sends Ollama's native `POST /api/pull` request to the selected endpoint. Therefore a download goes to the Ollama instance shown in the job message, and the Ollama service owns the model storage location. The exact path is controlled by Ollama's `OLLAMA_MODELS` setting; common Linux service/user locations are `/usr/share/ollama/.ollama/models` and `~/.ollama/models`. Check the Ollama service environment on the target host to determine the authoritative path. ## Security limits - Uploaded files are limited to 20 MiB each - Fetched URLs are limited to 15 MiB and a 30-second timeout - Private, loopback, link-local, reserved, multicast, and unspecified URL targets are blocked, including redirect destinations - Remote documents are inserted as untrusted content, not system instructions - Only locally installed models can be selected or loaded; chat does not download models ## Dependency The Hermes Plugin Section installs the declared `pypdf` dependency automatically before the plugin is committed. For a manual checkout, install it with: ```bash pip install -r requirements.txt ``` `pypdf` is required for text-based PDF extraction. Scanned/image-only PDFs need OCR and are not converted to text by this plugin.