A Proxmox VE cluster transforms individual hypervisor nodes into a unified, highly available virtualization platform. Clustering enables live migration of VMs between nodes, centralized management through any node's web interface, shared storage with Ceph, automatic VM restart on surviving nodes during hardware failures, and coordinated backup scheduling across the cluster. This guide covers enterprise cluster design, deployment, and configuration based on our production experience at Petronella Technology Group, Inc.
Cluster Planning
Node Count and Quorum
Proxmox clusters use a quorum-based voting system to prevent split-brain scenarios. Each node gets one vote, and the cluster requires a majority of votes to remain operational. This means a two-node cluster loses quorum if either node fails (not recommended without a QDevice). A three-node cluster tolerates one node failure. A five-node cluster tolerates two node failures. For production environments, we recommend a minimum of three nodes. This provides fault tolerance for a single node failure while keeping costs manageable.
If you need a two-node cluster for budget reasons, add a QDevice, an external witness that provides the tiebreaker vote. The QDevice is a separate host that runs the corosync-qnetd daemon, and the Proxmox documentation recommends it for two-node clusters.
The QDevice does not need to sit on the cluster LAN. Unlike Corosync itself, it connects to the cluster over TCP/IP, so it can run outside the cluster network and does not have to meet Corosync's low-latency requirement. The external host only needs network access to the cluster and the corosync-qnetd package. One QDevice host can serve several clusters.
Use a QDevice only with an even number of nodes. The Proxmox documentation supports it for even-sized clusters and discourages it for odd-sized ones. In an odd-sized cluster the QDevice provides N-1 votes instead of one, which turns the QDevice into a near single point of failure: if it goes down, the cluster cannot lose any other node without losing quorum.
Network Architecture
Enterprise Proxmox clusters should use dedicated networks for different traffic types. The cluster communication network (Corosync) carries heartbeat and cluster state information and should be an isolated, low-latency network. A dedicated VLAN or physical network is recommended, with redundant links for fault tolerance. The VM traffic network carries virtual machine network traffic and should have sufficient bandwidth for your workload requirements. The storage network carries Ceph replication traffic or iSCSI/NFS storage traffic and should be a high-bandwidth, low-latency network (10GbE or faster for Ceph). The management network provides web interface access and API communication.
At minimum, use two physical network interfaces per node: one dedicated to Corosync cluster traffic, and one for VM, storage, and management traffic. For production environments, four or more interfaces with bonding provide the best balance of performance and redundancy.
Corosync Network Requirements
Corosync needs consistent low latency, not bandwidth. The Proxmox documentation states that a dedicated 1 Gbit NIC is enough for cluster traffic in most situations, and it sets these requirements for every cluster:
- All nodes must reach each other on UDP ports 5405 to 5412 for Corosync, and on TCP port 22 for the SSH tunnel between nodes.
- Network latency between all nodes must stay under 5 milliseconds (LAN performance). The documentation warns that stable operation becomes unlikely with more than three nodes and latencies above roughly 10 ms, so keep each cluster inside one site.
- Date and time must be synchronized. Proxmox VE has used chrony as its default NTP daemon since version 7.
- All nodes should run the same Proxmox VE version.
- Online migration of virtual machines is only supported between nodes with CPUs from the same vendor.
Do not run Corosync and storage traffic on the same network. Ceph recovery traffic in particular can delay Corosync packets enough to cost the cluster its quorum. Corosync supports up to 8 links, so give it a second link on a different physical network. The cluster then keeps communicating when the dedicated network goes down.
Be careful with bonds under Corosync. The documentation advises against the balance-rr, balance-xor, balance-tlb, and balance-alb bond modes for Corosync traffic, and it strongly recommends setting bond-lacp-rate fast on both the node and the switch when you use LACP. A separate Corosync link on independent hardware is the safer design.
Corosync has used the Kronosnet transport with regular UDP unicast since Corosync 3.0, which arrived with Proxmox VE 6.0. Multicast support on your switches is no longer a requirement.
Storage Strategy
Choose your storage strategy based on your requirements. Local ZFS provides excellent performance with data protection through mirroring or RAIDZ. Each node manages its own storage, so VMs must be migrated to move between nodes (live migration works with local storage using Proxmox's storage replication feature). Ceph provides distributed, replicated storage accessible from all cluster nodes. VMs can live migrate between any nodes without storage movement because all nodes access the same Ceph pool. Ceph requires a minimum of three nodes and dedicated storage network bandwidth. Shared NFS or iSCSI provides a traditional shared storage model where an external storage appliance serves storage to all nodes.
Hardware Sizing for ZFS and Ceph
Plan the disk controllers before you order hardware. The Proxmox installation requirements state that neither ZFS nor Ceph is compatible with a hardware RAID controller, so both need disks presented directly to the host. The same page recommends SSDs with power-loss protection and discourages consumer SSDs.
Memory planning matters as much as disks. Proxmox recommends a minimum of 2 GB of memory for the operating system and Proxmox VE services, plus the memory you assign to guests, plus approximately 1 GB of memory for every TB of used storage under Ceph or ZFS. For Ceph, the current recommendation is at least 8 GiB of memory per OSD for good performance, and the OSD daemon requires 4 GiB by default. Leave headroom beyond that, because Ceph uses more memory during recovery, rebalancing, and backfilling than it does on a normal day.
For CPU, the rule of thumb in the Proxmox Ceph chapter is at least one CPU core or thread per Ceph service. A node that runs one monitor, one manager, and six OSDs should reserve 8 cores for Ceph alone.
For the Ceph network, Proxmox recommends at least 10 Gbps used exclusively for Ceph traffic. A single NVMe SSD can saturate 10 Gbps, so NVMe-backed clusters should plan for at least 25 Gbps. When in doubt, the documentation recommends three physically separate networks for high-performance setups: a 25 Gbps or faster network for Ceph internal cluster traffic, a 10 Gbps or faster network for Ceph public traffic, and a 1 Gbps network used only for Corosync.
Storage Replication for Local ZFS
Clusters that use local ZFS instead of Ceph can still protect guests with the Proxmox storage replication framework. It replicates guest volumes to another node using snapshots, so after the initial full sync it only sends changes. Replication runs every 15 minutes by default. You can set the interval anywhere from one minute to once a week, and you can rate-limit each job to protect the storage and the network.
Know the tradeoff before you pair replication with HA. The documentation allows the combination, but warns that you may lose the data written between the last sync and the moment the node failed. Workloads that cannot tolerate that gap belong on shared or distributed storage.
Cluster Deployment
Step 1: Install Proxmox VE on All Nodes
Install Proxmox VE on each node using the standard installation ISO. During installation, configure the management IP address and hostname for each node. Ensure all nodes can resolve each other by hostname (configure DNS or /etc/hosts entries). Verify that all nodes can reach each other on the cluster communication network.
Set the final hostname and IP configuration before you build the cluster. The Proxmox documentation is explicit that you cannot change either one after cluster creation. Join nodes while they are still empty: joining overwrites the existing configuration in /etc/pve, and a joining node cannot hold any guests because guest IDs could conflict with IDs that already exist in the cluster.
Step 2: Create the Cluster
On the first node, create the cluster through the web interface (Datacenter, Cluster, Create Cluster) or via the command line: pvecm create your-cluster-name. You can specify which network interface to use for cluster communication with the --link0 parameter, and add a redundant link with --link1.
Step 3: Join Additional Nodes
On the first node, open Datacenter, Cluster and copy the Join Information. On each additional node, paste that Join Information into the Join Cluster dialog, or run pvecm add with the address of an existing cluster node. The join process synchronizes cluster configuration, SSH keys, and certificate authority across all nodes.
After joining, verify the cluster status with pvecm status. All nodes should show as online, and the quorum should be met.
The full command-line sequence looks like this. Replace the names and addresses with your own values.
# on the first node
pvecm create CLUSTERNAME --link0 LOCAL-IP-ON-CLUSTER-NETWORK
# on each additional node
pvecm add IP-ADDRESS-CLUSTER --link0 LOCAL-IP-ADDRESS-LINK0
# verify from any node
pvecm status
pvecm nodes
You need the root password of an existing cluster node to add a new one. In the pvecm status output, confirm that the Quorate line reads Yes and that the expected vote count matches your node count.
Optional: Add a QDevice to a Two-Node Cluster
Install the corosync-qnetd package on the external witness host and the corosync-qdevice package on every cluster node, then run the setup command from one cluster node:
# on the external witness host
apt install corosync-qnetd
# on every cluster node
apt install corosync-qdevice
# on one cluster node
pvecm qdevice setup QDEVICE-IP
The setup copies the cluster SSH key to the witness host, so the root user on that host must allow key-based access, or password login for the duration of the setup. All cluster nodes must be online when you run it.
Step 4: Configure Ceph (Optional)
If you are using Ceph for distributed storage, install the Ceph packages on each node through the web interface (Datacenter, Ceph, Install). Configure the Ceph monitor and manager daemons on at least three nodes. Create OSDs (Object Storage Daemons) on each node's storage disks. Create a Ceph pool for VM storage. Verify Ceph health with ceph status, all placement groups should be active and clean.
Ceph configuration through the Proxmox web interface is straightforward and does not require manual editing of Ceph configuration files. The integration handles monitor configuration, OSD creation, pool management, and health monitoring.
The same steps are available from the command line through the pveceph tool:
pveceph install
pveceph init --network 10.10.10.0/24
pveceph mon create
pveceph mgr create
pveceph osd create /dev/sd[X]
pveceph pool create POOL-NAME --add_storages
Create three monitors. Proxmox states that you need at least 3 monitors for high availability and that small to medium-sized clusters do not need more than 3. At least one manager daemon is required as well. When you create a pool without options, Proxmox sets 128 placement groups, a size of 3 replicas, and a min_size of 2 replicas. Do not lower min_size to 1, which the documentation explicitly warns against.
Step 5: Configure High Availability
Proxmox HA automatically restarts tagged VMs on surviving nodes when a node fails. Configure HA by adding each VM as an HA resource, then use HA node affinity rules to define which nodes may host it. Current Proxmox VE releases migrate the older HA groups to these rules. Proxmox's HA manager uses fencing to ensure failed nodes are truly offline before restarting their VMs elsewhere, preventing the possibility of the same VM running on two nodes simultaneously.
Proxmox lists three requirements before you turn HA on: at least three cluster nodes for reliable quorum, shared storage for VMs and containers, and hardware redundancy everywhere. HA groups have been deprecated and migrated to HA node affinity rules since Proxmox VE 9.0, so build new clusters on rules from the start.
Adding a guest to HA takes one command per resource. The service ID is the resource type plus the VMID:
ha-manager add vm:100
ha-manager add vm:501 --state started --max_relocate 2
ha-manager status
An HA resource can only run on nodes where all of its dependencies exist. Keep HA guests on storage that every node can reach, on network interfaces that exist on every node, and free of device passthrough that only one node offers. If a dependency exists on a subset of nodes, restrict the resource to that subset with a strict node affinity rule.
Two services do the work. The cluster resource manager (pve-ha-crm) makes cluster-wide decisions, and the local resource manager (pve-ha-lrm) on each node carries them out. Fencing relies on a watchdog. A node that loses quorum cannot reset its watchdog, and if it holds active HA services it reboots when the watchdog times out after 60 seconds. Proxmox uses a hardware watchdog when one is configured and falls back to the Linux kernel softdog otherwise. Only after the failed node is fenced does the cluster recover its services on another node.
Decide on a shutdown policy before your first maintenance window. Proxmox ships four: Migrate, Failover, Freeze, and Conditional. Conditional remains the default for backward compatibility, and the documentation notes that some users may find Migrate behaves more as expected, because it live-migrates HA services away before the node shuts down. Set it under Datacenter, Options, HA Settings. For planned work on one node, use maintenance mode, which moves HA services away and returns them when you finish:
ha-manager crm-command node-maintenance enable NODENAME
ha-manager crm-command node-maintenance disable NODENAME
Step 6: Configure Backup
Set up Proxmox Backup Server (PBS) and configure backup jobs for your cluster. Our Proxmox Backup Server configuration guide covers the details. Create backup schedules that stagger across nodes to avoid overwhelming the backup storage. Configure retention policies based on your recovery point objectives. Test restore procedures to verify backup integrity.
Pick the backup mode on purpose. Proxmox offers three modes for VMs. Stop mode gives the highest consistency at the cost of a short downtime. Snapshot mode gives the lowest downtime with a small inconsistency risk. Suspend mode exists for compatibility, and the documentation recommends snapshot mode instead. Install and enable the QEMU guest agent in every VM: with the agent running, snapshot backups freeze and thaw the guest file system to improve consistency.
For the backup target, the Proxmox documentation recommends Proxmox Backup Server on a dedicated host, where it stores backups as de-duplicated chunks, and names an NFS server as a good alternative. On a slow backup target, enable backup fleecing. Fleecing caches old blocks in a local fleecing image instead of making guest writes wait on the backup target, which protects guest IO performance during the backup window at the cost of extra local storage.
Express retention as policy, not as a disk-full event. Backup jobs accept keep-last, keep-hourly, keep-daily, keep-weekly, keep-monthly, and keep-yearly options, and Proxmox processes them in that order. The pvescheduler daemon runs the jobs on a schedule format similar to systemd calendar events.
Enterprise Hardening
Firewall Configuration
Proxmox includes a built-in firewall that can be configured at the datacenter, node, and VM level. For enterprise deployments, enable the Proxmox firewall, create rules that restrict management access (port 8006) to authorized networks, allow cluster communication (Corosync) between nodes, allow Ceph traffic between nodes on the storage network, and restrict SSH access to authorized administrators.
The Proxmox firewall is completely disabled by default. You enable it cluster-wide with enable: 1 in the OPTIONS section of /etc/pve/firewall/cluster.fw, or under Datacenter, Firewall in the web interface. Once it is on, it blocks traffic to all hosts by default, with exceptions for traffic from the local network to the web interface, SSH, and other important services. Open an SSH session to a node before you enable it, so you still have a way in if a rule is wrong. Host-level rules in /etc/pve/nodes/NODENAME/host.fw take precedence over datacenter-level rules.
The default rule set shows which ports a cluster uses, and it is a useful checklist for upstream firewalls too:
- TCP 8006 for the web interface
- TCP 5900 to 5999 for VNC web console traffic
- TCP 3128 for SPICE proxy connections
- TCP 22 for SSH
- TCP 60000 to 60050 for live migration
- UDP 5405 to 5412 for Corosync on the cluster network
When the firewall is enabled, Proxmox generates the ACCEPT rules for Corosync automatically. Ceph ports are not opened for you, and the simplest way to allow them is the firewall macro that Proxmox provides for Ceph.
Authentication and Access Control
Proxmox supports multiple authentication backends including its built-in PVE authentication, LDAP, Active Directory, and OpenID Connect. For enterprise environments, integrate with your existing directory service. Configure role-based access control (RBAC) to limit user permissions to the minimum required for their responsibilities.
Require two-factor authentication for every administrator. A realm can enforce TOTP or YubiKey OTP for all of its users, and individual users can opt in to a second factor even when the realm does not enforce one. Keep in mind that the system root user can always log in through the Linux PAM realm, so protect that account accordingly.
Give automation its own credentials. API tokens allow stateless access to the REST API, and you can revoke a compromised token without disabling the user behind it. Leave privilege separation enabled so each token carries only the permissions you grant it. Use resource pools to group virtual machines, containers, and storage, then assign permissions to the pool instead of to every resource.
Certificate Management
Replace the default self-signed certificates with certificates from your internal CA or a public CA. Proxmox supports ACME (Let's Encrypt) for automated certificate management if your Proxmox nodes are accessible from the internet. For internal deployments, deploy certificates from your enterprise PKI.
Internal nodes are not locked out of ACME. The Proxmox ACME client supports both http-01 and dns-01 challenges. The dns-01 challenge proves domain ownership with a DNS record, so it works for nodes that the public internet cannot reach. You register one ACME account per cluster under Datacenter, ACME or with the pvenode command-line tool, and the documentation recommends the Let's Encrypt staging endpoint for first-time experiments because of rate limits.
Monitoring and Alerting
Proxmox's web interface provides basic monitoring of node and VM resources. For enterprise monitoring, integrate Proxmox with your monitoring stack. Proxmox VE can send metrics to external metric servers such as Graphite and InfluxDB. Configure alerts for node failures and HA events, storage capacity thresholds, Ceph health degradation, backup job failures, and resource utilization anomalies.
Our monitoring integration work pairs Proxmox with Prometheus, Grafana, and alerting, which gives operators real-time visibility into cluster health, storage utilization, and VM performance.
Proxmox VE currently supports three external metric server types: Graphite, InfluxDB, and OpenTelemetry. The definitions live in /etc/pve/status.cfg and you can edit them from the web interface. For alerts, the notification system routes events through matchers to targets. The built-in target types are Sendmail, SMTP, Gotify, and Webhook, so you can send a backup failure to a mailbox and a fencing event to a chat channel or a ticketing system.
Maintenance Procedures
Regular cluster maintenance includes applying Proxmox updates (one node at a time, migrating VMs before rebooting), monitoring Ceph health and rebalancing after hardware changes, verifying backup job success and testing restore procedures, reviewing firewall rules and access control configurations, and monitoring storage capacity and planning expansions.
Proxmox supports rolling updates across the cluster. Migrate VMs off a node, update and reboot it, verify it rejoins the cluster successfully, then proceed to the next node.
On each node, update with apt update followed by apt full-upgrade, or use the node's Updates panel in the web interface. The pve-enterprise repository is enabled by default and requires a valid subscription key, so assign subscriptions or configure your repositories before the first patch cycle. In a hyper-converged cluster, check the Ceph version requirement before every Proxmox VE major upgrade, because each major release requires a corresponding minimum Ceph major release.
Removing a Node Safely
Decommission nodes with care, because a mistake here can break the cluster. Move all virtual machines and containers away from the node first. If the node runs Ceph services, destroy its OSDs one at a time and wait for HEALTH_OK after each, then destroy its monitor and manager. Power the node off, and only then run pvecm delnode NODENAME from one of the remaining nodes. The Proxmox documentation calls it critical that the removed node never powers on again in the cluster network with its old configuration. Reinstall it before you reuse it.
After the removal, clean up what stays behind: the node's directory under /etc/pve/nodes, its public key in /etc/pve/priv/authorized_keys, and any HA rules that still reference the node.
Recovering Configuration Access Without Quorum
The cluster file system (pmxcfs) that holds /etc/pve becomes read-only when a node loses quorum. That is a safety feature, not a fault. If you must change configuration on a node without quorum and you understand the consequences, pvecm expected 1 sets the expected vote count to 1 and makes that node quorate. Use it to repair, never to keep running two halves of a partitioned cluster, because that is how split-brain damage happens.
Clusters That Host Regulated or AI Workloads
A virtualization cluster often ends up holding the most sensitive systems a company owns. If yours will host regulated data, fold the hardening steps above into your compliance program instead of treating them as a one-time build task. Our CMMC compliance and NIST 800-171 resources explain the requirements, and the cybersecurity and compliance glossary defines the terms your assessor will use.
The same cluster design also suits on-premise AI. Teams that want to keep model inference and training data off public clouds can pass physical GPUs through to guests on Proxmox nodes with PCI(e) passthrough. Our AI services hub and our comparison of private AI and cloud AI cover that side of the design.
Proxmox Cluster FAQ
How many nodes does a Proxmox cluster need?
Three is the practical minimum for production. Proxmox requires at least three nodes for reliable quorum and for high availability. A two-node cluster works when you add a QDevice as an external third vote.
What network latency does a Proxmox cluster require?
Under 5 milliseconds between all nodes, which the Proxmox documentation describes as LAN performance. It warns that stable operation is unlikely with more than three nodes once latency rises above roughly 10 ms.
Which ports does a Proxmox cluster use?
Corosync uses UDP 5405 to 5412 and the nodes use TCP 22 between each other. Administrators reach the web interface on TCP 8006. The default firewall rules also cover TCP 5900 to 5999 for the VNC console, TCP 3128 for SPICE, and TCP 60000 to 60050 for live migration.
Do I need Ceph to run Proxmox high availability?
No, but you need shared storage that every node can reach. Ceph is one option, and NFS or iSCSI from an external appliance is another. Local ZFS with storage replication also works with HA, with the documented risk of losing data written since the last replication run.
Can I use a hardware RAID controller with Ceph or ZFS?
No. The Proxmox installation requirements state that neither ZFS nor Ceph is compatible with a hardware RAID controller. Present the disks directly to the host.
How does Proxmox prevent the same VM from running on two nodes?
Through fencing. A node that loses quorum cannot reset its watchdog and reboots after the 60 second watchdog timeout if it holds active HA services. The cluster recovers those services on another node only after the failed node is fenced.
How do I patch a Proxmox cluster without downtime?
One node at a time. Put the node into HA maintenance mode or live-migrate its guests away, run apt update and apt full-upgrade, reboot, confirm the node rejoins with pvecm status, and then move to the next node.
Getting Enterprise Support
For organizations deploying Proxmox clusters in production, Petronella Technology Group, Inc. provides cluster design and deployment services, Ceph storage architecture and tuning, HA configuration and failover testing, monitoring integration and dashboard setup, and ongoing managed support and maintenance. Our team runs Proxmox in production and brings that operational experience to client engagements.
Plan Your Proxmox Cluster With Us
Petronella Technology Group, Inc. has provided cybersecurity, compliance, and managed IT services from Raleigh, NC since 2002. Call Penny at 919-348-4912 or contact us to talk through your cluster design.
Related Resources
- Cloud Repatriation Services
- VMware to Proxmox Migration
- Managed IT Services Guide
- VMware to Proxmox Migration Guide
A decision tree, a workload placement worksheet and a hardening checklist for Proxmox VE 9.2 and Docker. Free PDF.
Get the free guide