Virtualization Migration Should Account for Operational Continuity and AI Scalability
IT DAILY ·
✦ AI Summary
Following Broadcom's policy overhaul on licensing, companies and public-sector agencies have begun in earnest to review alternatives to VMware. Evaluation criteria are shifting away from whether to replace the hypervisor and toward operational continuity, AI infrastructure scalability, and whether mission-critical workloads can be migrated without disruption. Seo Young-seok, vice president at Orquestro, said VM migration and actual operational continuity are different matters, and stressed the need for workflow verification that can identify root causes and support responses during failures, as well as pre-migration diagnosis and rollback for closed-network environments.
Following Broadcom's policy overhaul on licensing, companies and public-sector agencies have begun in earnest to review alternatives to VMware. As a result, the evaluation criteria for virtualization migration are also shifting away from simply whether to replace the hypervisor, and toward operational continuity and the scalability of AI infrastructure, with AI infrastructure scalability emerging as the true yardstick for virtualization migration.
Amid this shift, whether mission-critical workloads can be migrated without disruption is being cited as a key consideration at the implementation stage. Another major question is whether the existing operations team can immediately identify the cause of a failure and continue responding.
Seo Young-seok, vice president at Orquestro, said in a recent in-house webinar that VM migration and the uninterrupted operation of actual infrastructure are challenges on different levels. He said simple side-by-side comparisons of feature checklists should be avoided, and that practical workflow verification is needed to see whether operators can determine their next action when a failure occurs.
The photo shows a capture from an Orquestro webinar.
It was also noted that the virtualization transition requires continuity with next-generation AI and container infrastructure.
Seo Young-seok said data center environments undergoing infrastructure modernization are evolving from physical-server-centered setups, through a mixed stage of server VMs, containers and multicloud, toward an architecture that accommodates GPU and AI workloads.
He added that even if operations are currently centered on VMs, containers, GPUs and AI services will soon be integrated within the same data center. He also said that if resources are operated separately as fragmented point solutions, the resource allocation framework and failure response framework become disconnected, and operational complexity increases exponentially under separate operations.
He also pointed out that if infrastructure is built with too much focus on replacing virtualization functions alone, there is a risk that the architecture will have to be completely redesigned later when Kubernetes and AI services need to be expanded.
It was pointed out that migration risk is high in closed-network environments that require a high level of security, such as in the public sector and financial industry. Seo said that during hypervisor transitions, simple data copying alone makes it difficult to guarantee normal booting, network connectivity and application integrity.
Seo explained that failure rates rise sharply in closed networks with significant restrictions on network interconnection. Accordingly, he said pre-migration diagnosis and rollback procedures based on verified solutions are essential.
As a concrete verification criterion for ensuring operational stability on site, the ability to support operational decision-making was presented. Seo raised five practical questions that operators face during failures and maintenance, rather than focusing on the number of dashboard screens.
A representative case involved an unexplained performance slowdown. In that case, there was room in the metrics, with CPU utilization at 60% and memory utilization at 72%, but users complained of service delays.
Seo said simple average utilization rates have limitations when it comes to explaining where system tasks are waiting. He added that it is necessary to determine whether run-queue delays are occurring based on vCPU, and whether major page faults or contention are occurring during the memory allocation process. He also said there is a need to provide deep metrics intuitively so that the pool of possible causes can be narrowed down.
He also said automation is needed for routine inspection procedures such as host firmware updates. Before inspection, it is necessary to pre-verify the target host's spare capacity, and to pre-check whether deployment policies such as anti-affinity are being violated. He added that if constraints arise, the work should not be forced through but safely halted, and that a system for reporting the reason for the halt is also needed.
Seo said that if there are no abnormalities in the virtual layer, rack configuration mapping is also needed. Through this, he said it should be possible to trace physical NIC errors on the host and port errors on the upper ToR physical switch. That, he explained, would make it possible to identify the root cause without back-and-forth between infrastructure teams.
Orquestro suggested that it is necessary to establish an operational framework that connects failure response and operations management in five stages, from detection to evidence. Stage 1 is 'detection,' or identifying signs of abnormalities. Stage 2 is 'correlation analysis,' or connecting alarms, resources and task history into a single timeline. Stage 3 is 'physical tracing,' or tracing to the physical network switch. Stage 4 is 'action,' or carrying out safe measures based on pre-verified acceptability. Stage 5 is 'evidence,' or documenting what changed before and after the failure.
Seo Young-seok then said that service recovery alone is not enough as the standard for operational completeness. He explained that operations are complete only when audit logs and recordkeeping systems can show what changed and what actions were taken. He also said the evaluation criteria for virtualization alternatives need to change, stressing that if the old criteria were geared toward replicating the feature lists of foreign solutions, the post-transition criteria should be whether the operations team can actually control and resolve complex failure situations.
Source: IT DAILY · Kwon Young-seok
Original: https://www.itdaily.kr/news/articleView.html?idxno=242032
References
This article was produced with the help of an automated content generation algorithm.
Source: IT DAILY
View originalThis article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.