All articles

Thousands of VMs on Proxmox: What Squark's VMware Exit Shows

Squark's Proxmox platform handles thousands of VMs across several regions. The useful lesson is how repeatable migration and operating decisions scale.

Computer Port IT Solutions5 min read
Long rows of enterprise server racks in a large data centre with orange status lights

A virtualization platform handling a few dozen workloads can hide design problems for a long time. At provider scale, those problems become operating problems very quickly.

That is what makes Proxmox's published Squark case worth reading. Squark is a French infrastructure-as-a-service provider whose virtualization environment handles thousands of virtual machines across several regions and availability zones. According to the official story, each zone corresponds to 2,000 compute cores and 30 TB of RAM.

The company began moving its IaaS platform away from VMware vSphere in 2015. Proxmox gradually replaced the VMware suite, and the complete transition took place in 2016.

This is not a story about a single maintenance window. It is a story about turning a platform change into a repeatable operating process.

The number matters, but the method matters more

Thousands of VMs make a strong headline. The more useful detail is how Squark approached the change.

The official case study says the team developed an internal workflow to manage and schedule virtual-to-virtual migrations because the import tooling available today did not exist at the time. It also benchmarked the HPE Synergy hardware used for the target platform before broader production deployment.

Those two decisions are easy to overlook:

  • make the migration repeatable before increasing volume
  • validate the target hardware before production depends on it

At scale, the migration method needs to survive more than one engineer, one cluster or one weekend. Each wave should have the same entry checks, conversion steps, validation evidence and rollback decision.

A large VMware exit is an operations project

The hypervisor is only one part of the target state. Squark's platform spans several regions and availability zones, so the operating model also needs to cover:

  • cluster standards and version ownership
  • capacity and failure-domain planning
  • network and storage design
  • workload placement
  • backup and restore evidence
  • maintenance and update procedures
  • support and escalation
  • visibility across multiple locations

A platform can be technically capable and still be difficult to operate if these decisions are left to each cluster owner.

This is why a VMware-to-Proxmox assessment should not end with a VM count. It should show how the environment will be built, supported and reviewed after the last migration wave is complete.

Migration waves need real acceptance criteria

Proxmox now includes an Import Wizard for direct guest import from other hypervisors. That reduces manual conversion work, but it does not decide whether a workload is ready.

Before a VM enters a migration wave, the team still needs to know:

  1. Who owns the application and can approve a test?
  2. Which network, storage, identity and licensing dependencies follow it?
  3. Is there current backup and restore evidence?
  4. Does the target platform have enough capacity during maintenance or failure?
  5. What proves that the migrated workload is healthy?
  6. What condition triggers rollback?
  7. Who owns monitoring, backup and support after cutover?

The first workloads should test the process, not only the import tool. A representative pilot is more valuable than choosing only the easiest VM and assuming the rest will behave the same way.

Hardware validation is part of the migration plan

Squark's team ran benchmarks on its HPE Synergy 480 Gen10 environment before expanding production use. That is a useful reminder that compatibility is not the same as suitability.

A target host may boot and recognize its hardware while still leaving open questions about:

  • network throughput and redundancy
  • storage latency under realistic load
  • controller and firmware behavior
  • failure and recovery time
  • live migration performance
  • capacity when a node is unavailable
  • operational access during an incident

The target design should be tested with the same constraints that production will face. A quiet lab does not reveal every problem that appears when backup, replication, maintenance and customer workloads compete for the same resources.

What should an enterprise take from this case?

Squark's story shows that Proxmox can sit behind a large, multi-region service-provider environment. It does not mean every VMware estate should copy the same architecture.

The practical lesson is to build a migration system that can scale:

  • start with an accurate estate and dependency inventory
  • define the target operating model before the first wave
  • benchmark the actual hardware and storage path
  • prove backup and restore while both platforms still exist
  • use repeatable wave criteria and validation evidence
  • assign day-two ownership for clusters, backup, access and support

Scale magnifies unclear decisions. It also rewards a migration process that can be repeated, reviewed and improved.

Computer Port IT Solutions is a Proxmox Silver Partner in India. We help organizations assess VMware estates, design Proxmox VE and Proxmox Backup Server environments, validate representative workloads and plan controlled migration waves.

Explore our VMware-to-Proxmox services:

https://computerport.in/vmware-to-proxmox

Primary source

Computer Port has no relationship with Squark. This article analyses the public Proxmox customer story and does not imply otherwise.

ProxmoxVMware MigrationVirtualizationCloud InfrastructureData Center