# Hermes Ollama Models Native-like Hermes dashboard plugin for local Ollama model management and chat. ## Chat capabilities - Select an installed Ollama model and load it into memory - Chat through Ollama's native `/api/chat` endpoint - Attach screenshots and JPEG/PNG/WebP images for vision-capable models - Attach text PDFs; PDF text is extracted with `pypdf` - Add public HTTP/HTTPS URLs for HTML/text, images, or PDFs - View live host RAM and swap statistics - View Ollama's loaded-model memory split: total, GPU VRAM, and normal RAM/offload - View NVIDIA GPU telemetry when `nvidia-smi` is available - Chat is the default view when the Ollama Models plugin opens - Natural composer behavior: Enter sends; Shift+Enter creates a new line - Paste images directly into the composer and drag/drop images, PDFs, and text files - Streamed Ollama responses with a real Stop action that cancels the active request - Minimized-by-default expandable thinking/progress details with live stage, elapsed time, event, and character counters The chat transcript and selected model are persisted in the Hermes server's SQLite database at `~/.hermes/ollama-manager/chat.sqlite3`, so conversations can be listed and resumed from another browser or after a dashboard restart. Use **New conversation** to start a separate thread and **Clear chat** to delete the selected shared conversation. Uploaded files remain temporary; attachment names/types/URLs are retained as metadata, not raw file contents. ## Performance metrics Each completed request records Ollama-provided counters and timing when available: - Time to first token (TTFT) - Total request latency - Prompt/input token count and prompt tokens/sec - Output token count and output tokens/sec - Total tokens - Ollama total/load/prompt-evaluation/generation durations - Model, request status, request ID, and error details The dashboard shows per-conversation metrics and rolling aggregates across the latest 100 requests. Values are shown as unavailable when Ollama does not provide them; the plugin does not estimate token counts. ## Catalog and download filters The Popular and Available downloads views only show models with known size and expected RAM estimates that fit the running host's installed RAM, detected from `/proc/meminfo` and reported in the API filter metadata. The limit is calculated at runtime, so it follows RAM upgrades or downgrades without a hard-coded threshold. Available downloads also includes known fitting variants from installed model families. The Available downloads controls support: - Dense or MoE model type - Abilities such as completion, thinking, tools, vision, audio, and video when supplied by the catalog - Organization by upstream popularity, newest (`modified_at`), smallest size, largest size, or name - Search across model names, families, strengths, and capabilities Popularity and date ordering use upstream metadata only; the plugin does not invent popularity, dates, RAM requirements, or token metrics. ## Installation This plugin is installable from the Hermes dashboard Plugin Section using the repository URL: ```text https://gitea.beyondcloud.solutions/dennii/Hermes-Ollama_Models.git ``` In Hermes Dashboard, open **Plugins**, choose **Install from repository**, enter the URL above, and install. The repository contains the root `plugin.yaml`, dashboard manifest, backend API, compiled frontend bundle, stylesheet, and an opt-in prerequisite declaration. On Linux, the Hermes installer will verify Ollama, install it with the official Ollama installer when missing, and install the plugin's `pypdf` dependency before committing the plugin into `~/.hermes/plugins/`. Ollama installation requires the Hermes container to run as root, which is the expected configuration for a privileged ZimaOS deployment. If the container is not running as root, the plugin install stops without enabling a partially configured plugin. After installation or an update, restart only the Hermes dashboard service if requested by the installer. The dashboard now includes a connection panel in the header. Enter an Ollama base URL, choose **Local** or **Remote**, and use **Test** before **Save**. Saved URLs are stored in Hermes plugin data with file permissions restricted to the Hermes account. Credential-bearing URLs are rejected. The plugin probes the configured local endpoint and known Docker endpoints, and tests a configured remote endpoint. A successful connection reports the Ollama version and installed model count. The active local endpoint continues to drive chat, status, load, and runtime telemetry. A URL saved from the dashboard takes precedence over the initial `OLLAMA_HOST` environment fallback, so the field can correct or replace an inherited container setting. When both distinct local and remote endpoints are online, downloading or re-downloading a model opens a destination chooser. The selected target is recorded on the job and the pull is sent to that endpoint. Remote downloads do not alter the local installed-model inventory. The connection rows display the actual model names returned by each endpoint. If Local and Remote contain the same URL, the dashboard marks the duplicate and treats it as one download target; enter the actual second Ollama server URL to enable independent remote downloads. The plugin supports both deployment types: - **Physical/Linux Hermes:** Ollama defaults to `http://localhost:11434`. If Ollama is missing, the prerequisite installer can install it on a root-run host. - **Docker Hermes:** Ollama should normally run as a separate service or on the Docker host. The plugin does not install Ollama inside the Hermes container. It first probes `http://ollama:11434` and then `http://host.docker.internal:11434`. Set `OLLAMA_HOST` to the reachable Ollama base URL, without `/v1`, for example `http://ollama:11434`, `http://host.docker.internal:11434`, or the host address shown by your installation, such as `http://:11434`. For host access on Linux Docker, add this to the Hermes service when needed: ```yaml extra_hosts: - "host.docker.internal:host-gateway" environment: OLLAMA_HOST: http://host.docker.internal:11434 ``` For Compose service-to-service networking, use the Ollama service name instead: ```yaml environment: OLLAMA_HOST: http://ollama:11434 ``` The dashboard plugin API is mounted when the dashboard starts. Restart Hermes after installing or enabling the plugin; otherwise the plugin page can load while `/api/plugins/ollama-manager/status` still returns 404. ## Download storage The plugin does not store model blobs in Hermes. It sends Ollama's native `POST /api/pull` request to the selected endpoint. Therefore a download goes to the Ollama instance shown in the job message, and the Ollama service owns the model storage location. The exact path is controlled by Ollama's `OLLAMA_MODELS` setting; common Linux service/user locations are `/usr/share/ollama/.ollama/models` and `~/.ollama/models`. Check the Ollama service environment on the target host to determine the authoritative path. ## Security limits - Uploaded files are limited to 20 MiB each - Fetched URLs are limited to 15 MiB and a 30-second timeout - Private, loopback, link-local, reserved, multicast, and unspecified URL targets are blocked, including redirect destinations - Remote documents are inserted as untrusted content, not system instructions - Only locally installed models can be selected or loaded; chat does not download models ## Dependency The Hermes Plugin Section installs the declared `pypdf` dependency automatically before the plugin is committed. For a manual checkout, install it with: ```bash pip install -r requirements.txt ``` `pypdf` is required for text-based PDF extraction. Scanned/image-only PDFs need OCR and are not converted to text by this plugin.