Skip to main content
This guide provides the recommended server resources for an on-premise installation of VARIOS AI. The language models run at your model provider or on a separate model server; VARIOS AI itself does not require a GPU. The recommendations are grouped into five tiers XS, S, M, L, and XL. Without users, VARIOS AI occupies about 5 GB of memory and 15 GB of disk space.
The values are reference values for planning. Review utilization after rollout and adjust the tier if necessary.

Recommendation by headcount

Count all people with access to VARIOS AI. It is assumed that on a working day 40% of them use VARIOS AI and send an average of 8 messages. Load peaks are taken into account.
If usage in your organization is mandatory or very widespread, choose the next higher tier.

Recommendation by concurrently active users

Use this table if you know the expected concurrency, for example from a pilot phase. A concurrently active user is currently chatting and sends a message about every two to three minutes. Open browser tabs without interaction do not count. If you measured concurrency at normal times, use three times the value as a reserve for peak times before reading the tier. If the two tables yield different tiers, the higher one applies.

Storage during operation

The database grows with the chats. With mixed usage of text and documents, this averages about 100 KB per question, or about 80 GB per year per 1,000 employees. The exact value depends heavily on usage; pure text chats need considerably less. With a chat retention period of 90 days, it stays at about 20 GB per 1,000 employees. Uploaded files add to this at their full size; their volume depends on usage and should be planned generously. Knowledge bases require about 7 MB per 100 pages of document text. Plan twice the database size for database backups.

What shifts the values

  • Knowledge bases and file uploads. Large volumes of documents increase the load of the DLP analysis. Load large knowledge bases outside working hours or choose one tier higher. For special requirements, such as DLP scanning of very many or very large documents, the DLP analysis can be moved to a dedicated server. Please contact support for this.
  • API integrations. Applications that send requests via API keys do so without pauses and often several requests in parallel. Count each concurrently running API request as three concurrently active users.

Distributed installation

For more than 2,000 employees or 200 concurrently active users, VARIOS AI is distributed across several servers. We plan this installation together with you; please contact support.

High availability

VARIOS AI distributes load across several servers when needed, but currently does not offer clustering with automatic failover. For high availability, we recommend running it on a virtualization platform with its own HA features, such as VMware vSphere HA, Proxmox HA, or a Hyper-V failover cluster. If a host fails, the platform restarts the virtual machine on another host. Complement this with regular backups of the database and the data directory.