Hermes Ollama Models

Native-like Hermes dashboard plugin for local Ollama model management and chat.

Chat capabilities

  • Select an installed Ollama model and load it into memory
  • Chat through Ollama's native /api/chat endpoint
  • Attach screenshots and JPEG/PNG/WebP images for vision-capable models
  • Attach text PDFs; PDF text is extracted with pypdf
  • Add public HTTP/HTTPS URLs for HTML/text, images, or PDFs
  • View live host RAM and swap statistics
  • View Ollama's loaded-model memory split: total, GPU VRAM, and normal RAM/offload
  • View NVIDIA GPU telemetry when nvidia-smi is available
  • Chat is the default view when the Ollama Models plugin opens
  • Natural composer behavior: Enter sends; Shift+Enter creates a new line
  • Paste images directly into the composer and drag/drop images, PDFs, and text files
  • Streamed Ollama responses with a real Stop action that cancels the active request
  • Minimized-by-default expandable thinking/progress details with live stage, elapsed time, event, and character counters

The chat transcript and selected model are persisted in the Hermes server's SQLite database at ~/.hermes/ollama-manager/chat.sqlite3, so conversations can be listed and resumed from another browser or after a dashboard restart. Use New conversation to start a separate thread and Clear chat to delete the selected shared conversation. Uploaded files remain temporary; attachment names/types/URLs are retained as metadata, not raw file contents.

Performance metrics

Each completed request records Ollama-provided counters and timing when available:

  • Time to first token (TTFT)
  • Total request latency
  • Prompt/input token count and prompt tokens/sec
  • Output token count and output tokens/sec
  • Total tokens
  • Ollama total/load/prompt-evaluation/generation durations
  • Model, request status, request ID, and error details

The dashboard shows per-conversation metrics and rolling aggregates across the latest 100 requests. Values are shown as unavailable when Ollama does not provide them; the plugin does not estimate token counts.

Catalog and download filters

The Popular and Available downloads views only show models with known size and expected RAM estimates that fit the running host's installed RAM, detected from /proc/meminfo and reported in the API filter metadata. Available downloads also includes known fitting variants from installed model families. The dashboard can filter downloads by Dense/MoE type or ability (completion, thinking, tools, vision, audio, or video), and organize them by popularity, newest, size, or name.

Security limits

  • Uploaded files are limited to 20 MiB each
  • Fetched URLs are limited to 15 MiB and a 30-second timeout
  • Private, loopback, link-local, reserved, multicast, and unspecified URL targets are blocked, including redirect destinations
  • Remote documents are inserted as untrusted content, not system instructions
  • Only locally installed models can be selected or loaded; chat does not download models

Dependency

Install the plugin's Python dependency in the Hermes runtime environment:

pip install -r requirements.txt

pypdf is required for text-based PDF extraction. Scanned/image-only PDFs need OCR and are not converted to text by this plugin.

S
Description
No description provided
Readme
1.1 MiB
Languages
Python 97.6%
Shell 2.4%