Upgrading NSX Manager in a Federated VCF Environment

Upgrading NSX Manager in a Federated VCF Environment | Farrukh's Tech Blog
VMware VCF · NSX Federation · Deep Dive

Upgrading NSX Manager
in a Federated VCF Environment

A step-by-step architect's guide for upgrading NSX 4.1.2.3 → 4.2.3.1 when SDDC Manager has no visibility of Global Managers — and why sequence is everything.

● VCF 5.x ● NSX Federation 4.1.2.3 → 4.2.3.1 April 2025

Upgrading NSX in a standard VCF workload domain is a well-understood workflow — SDDC Manager owns the lifecycle, orchestrates the upgrade bundle, and walks you through a pre-check → upgrade → validation loop. But introduce NSX Federation — with its Global Manager / Local Manager topology — and that comfortable automation suddenly has a blind spot: SDDC Manager has no visibility of Global Managers whatsoever.

Get the sequence wrong, and you can end up with a Local Manager running a newer NSX version than your Global Manager. Federation's N±1 interoperability rule means that is a hard-stop condition. This post walks through the complete, architect-level upgrade sequence for moving from NSX 4.1.2.3 → 4.2.3.1 in a federated VCF environment.

Section 01

Understanding the Architectural Blind Spot

Before any upgrade activity, you must understand what SDDC Manager sees and what it doesn't.

NSX Federation — Multi-site Architecture with Global Manager and Local Managers
NSX Federation Architecture — Global Manager and Local Managers
NSX-T / NSX Federation topology — Global Manager (Active/Standby) and per-site Local Managers

In a federated NSX deployment inside VCF:

  • Local Managers (LM) are registered as part of VCF workload domains. SDDC Manager sees them, manages their lifecycle, and upgrades them.
  • Global Managers (GM) are deployed independently and registered to SDDC Manager's inventory only as an external reference — SDDC Manager cannot upgrade them.
  • This means Global Manager upgrade is entirely manual, and must always happen before the Local Manager upgrade is triggered via SDDC Manager.
⚑ Fact-Checked — The "LM must never lead GM" rule has changed

A common misconception (including in earlier drafts of this post) is that allowing LM to exceed the GM version will categorically break federation. This was true prior to NSX 4.1.1, but no longer applies to 4.1.1+ or 4.2. Starting with NSX 4.1.1, and explicitly confirmed for NSX 4.2, upgrades can occur in any order — LM first or GM first — and federation sync is maintained across any version combination between 3.2 and 4.2.

That said, for VCF deployments on 4.1.x → 4.2, Broadcom's documented procedure still prescribes upgrading GM manually first for two specific reasons: (1) VCF BOM and SDDC Manager orchestration alignment, and (2) a resolved defect in 4.1.x that required all sites to be upgraded before moving to NSX 4.2. GM-first is still the right operational call — just understand why, so you're not cargo-culting an outdated rule.

📄 Broadcom TechDocs — Upgrading NSX Federation (NSX 4.2)
📄 Broadcom TechDocs — Upgrading NSX Federation (NSX 4.1)

Section 02

The Real Interoperability Model — What Changed in 4.1.1 and 4.2

The old N±1 rule — where GM had to be upgraded before LM at all times — applied only up to NSX 4.1.0. Broadcom fundamentally relaxed this constraint in subsequent releases. Understanding the version-specific rules is essential before planning your sequence.

Global Manager
4.2.3.1
⇌
Local Manager
4.1.2.3
✓ Actual Compatibility Rules — NSX 4.1.1+ and 4.2
  • Any upgrade order is supported (LM-first or GM-first) ✓
  • GM and LM sync is maintained across any version combination 3.2–4.2 ✓
  • Old N±1 rule → Applies only to NSX 4.1.0 and earlier
  • VCF procedure still prescribes GM-first for BOM + defect reasons

During the interim window — after you've upgraded GM to 4.2.3.1 but before SDDC Manager upgrades LM — your environment sits in a mixed-version state (GM 4.2.3.1, LM 4.1.2.3). Per Broadcom's documentation, federation sync continues uninterrupted in this state. The GM-first sequence is followed here because the official VCF 5.x upgrade procedure mandates it, not because the architecture requires it.

Section 03

Phase 1 — Pre-Upgrade Validation

Before touching a single component, perform thorough environmental health checks. Upgrades that fail mid-way in federated environments are significantly harder to recover from than in standalone deployments.

1.1 — Federation Health

1

Validate GM ↔ LM Channel Status

In Global Manager UI → System → Location Manager — all sites must show ACTIVE. Any DEGRADED or STANDBY alarm must be resolved before proceeding.

2

Check Config Replication Sync State

Verify no pending replication lag from GM to LM. Push a test config change and confirm propagation before upgrade.

3

Review Broadcom Interoperability Matrix

Confirm vCenter, ESXi, and vSAN versions in the target workload domain are all compatible with NSX 4.2.3.1. Use the VMware Interoperability Matrix at interopmatrix.vmware.com.

4

Backup GM and LM (All Nodes)

Trigger a manual NSX configuration backup for both Global Manager and all Local Managers via System → Backup & Restore. Confirm backup file is written and accessible.

5

Confirm VCF BOM Alignment

In SDDC Manager, validate that the VCF release bundle you are upgrading to includes NSX 4.2.3.1 in its Bill of Materials. SDDC Manager will not offer an NSX version that isn't in its BOM.

6

Confirm No Active Span Operations

Ensure no stretched segment migrations, HCX workload moves, or cross-site DR operations are in-flight. Pause or complete these before upgrade windows open.

Section 04

Phase 2 — Upgrade Global Manager (Manual)

ℹ SDDC Manager is not involved here

This entire phase is performed directly in the NSX Global Manager UI or via NSX API. SDDC Manager has zero visibility of this operation. You must complete this phase yourself before triggering anything via SDDC Manager.

4.1 — Active / Standby GM Pair

NSX Global Manager · System · Lifecycle Management · Upgrade
NSX Upgrade Coordinator — Global Manager
NSX Global Manager · Upgrade Coordinator
NSX Global Manager — Upgrade Coordinator showing bundle upload and pre-check phase
Upgrade Sequence — Active/Standby GM Pair
# Step 1 — Upload upgrade bundle to STANDBY Global Manager
Action : System → Lifecycle Mgmt → Upgrade
Upload : VMware-NSX-4.2.3.1-upgrade-bundle.mub
Target : Standby GM only

# Step 2 — Run pre-check on Standby GM
Action : Run Prechecks → Resolve all WARNINGs/ERRORs

# Step 3 — Execute upgrade on Standby GM
Action : Start Upgrade → Monitor until 100% complete
Validate: Standby GM reports healthy, reachable, version = 4.2.3.1

# Step 4 — Promote Standby to Active (planned failover)
Action : System → Location Manager → Promote Standby GM to Active
Confirm : New Active GM = 4.2.3.1 | Old Active GM now = Standby (4.1.2.3)

# Step 5 — Upgrade the original Active (now Standby) GM
Action : Repeat upgrade on remaining node
Validate: Both GMs = 4.2.3.1 | Active/Standby replication healthy

# Step 6 — Confirm all federation channels
Check  : System → Location Manager → all sites ACTIVE
Interim: GM = 4.2.3.1 | LM = 4.1.2.3 → N±1 valid, proceed
    
ℹ Single Active GM

If your environment has only a single Active GM (no standby pair), simply upload the bundle, run pre-checks, and execute the upgrade directly. There is no failover step. The GM will be unavailable for the duration of its upgrade — plan your change window accordingly, as no cross-site config pushes can occur during this window.

Section 05

Phase 3 — Upgrade Local Managers via SDDC Manager

Now that Global Manager is on 4.2.3.1 and federation channels are confirmed healthy, SDDC Manager can safely orchestrate the Local Manager upgrade. This is where the standard VCF lifecycle management workflow takes over.

SDDC Manager · Lifecycle Management · Upgrade
SDDC Manager Lifecycle Management
SDDC Manager · Lifecycle Management · Upgrade Workflow
SDDC Manager — Lifecycle Management upgrade workflow showing NSX (Local Manager) as a component target
1

Download the VCF Release Bundle

SDDC Manager → Lifecycle Management → Bundle Management. Download the target VCF bundle containing NSX 4.2.3.1 in its BOM. Confirm bundle is in AVAILABLE state.

2

Initiate Workload Domain Upgrade

Navigate to Lifecycle Management → Upgrade → select the target Workload Domain. SDDC Manager presents the component upgrade order: NSX → vCenter → ESXi/vSAN.

3

Run Pre-Checks — Resolve All Issues

SDDC Manager will run environment pre-checks. Do not proceed with any WARNING or ERROR state. Common blockers: certificate expiry, vSAN health failures, ESXi host connectivity issues.

4

Execute NSX Local Manager Upgrade

SDDC Manager upgrades the 3-node LM cluster in a rolling fashion (node-by-node). Monitor via both SDDC Manager UI and the NSX Manager UI simultaneously for any anomalies.

5

NSX Edge Cluster Upgrade (Automatic)

SDDC Manager orchestrates Edge node upgrades as part of the NSX lifecycle step. Edge nodes go one-by-one with traffic continuity maintained via BFD/ECMP failover on the T0 gateway.

6

vCenter and ESXi/vSAN Upgrades

SDDC Manager continues with vCenter (if in BOM), then ESXi cluster-by-cluster. Host upgrades use vSphere DRS-based DPM evacuation — confirm DRS is enabled and automation level is set appropriately.

NSX Manager · Upgrade Coordinator · Pre-Checks
NSX Upgrade Pre-check Dashboard
NSX Upgrade Coordinator · Pre-Check Results
NSX Upgrade Coordinator — Pre-check results showing component health before upgrade execution
Section 06

Phase 4 — Post-Upgrade Federation Validation

Both GM and LM are now on 4.2.3.1. Do not close your change window until all of the following validation points have been confirmed.

Post-Upgrade Validation Checklist
# 1. Federation Channel Health
Location : GM UI → System → Location Manager
Expected : All sites = ACTIVE  |  No DEGRADED / PARTIAL sites

# 2. Config Sync Validation
Action   : Push a test config change (e.g., tag on a segment) from GM
Expected : Change propagates to LM within expected replication window

# 3. Stretched Segment / Gateway Policy
Location : GM UI → Networking → Segments / Gateway Policies
Expected : No objects in PARTIAL_SUCCESS or ERROR realisation state

# 4. BGP / Routing Table Validation
Action   : SSH to T0 SR Edge nodes at each site
Command  : get logical-router <UUID> bgp neighbor summary
Expected : All BGP sessions ESTABLISHED | route counts stable

# 5. NSX Edge Cluster Health
Location : LM UI → System → Fabric → Nodes → Edge Transport Nodes
Expected : All Edge nodes = UP | Deployment status = NODE_READY

# 6. Alarm Review
Location : GM UI and LM UI → Alarms
Expected : No new CRITICAL or HIGH alarms post-upgrade

# 7. Datapath Verification (Optional but Recommended)
Action   : Run a cross-site ping/traceroute between stretched segment VMs
Expected : Traffic flows correctly across federation sites
    

Summary

Complete Upgrade Sequence at a Glance

Step Action Executed By Tool
01 Backup GM + LM (all nodes) Manual NSX UI / API
02 Validate federation health (all sites ACTIVE) Manual NSX Global Manager UI
03 Confirm VCF BOM includes NSX 4.2.3.1 SDDC Mgr SDDC Manager UI
04 Upgrade Standby GM → failover → upgrade original Active GM Manual NSX Global Manager UI
05 Validate GM health + federation channels (ACTIVE) Manual NSX Global Manager UI
06 Trigger Workload Domain upgrade via SDDC Manager SDDC Mgr SDDC Manager UI
07 SDDC Manager upgrades NSX Local Manager (rolling) SDDC Mgr SDDC Manager UI
08 SDDC Manager upgrades NSX Edge cluster SDDC Mgr SDDC Manager UI
09 SDDC Manager upgrades vCenter + ESXi/vSAN SDDC Mgr SDDC Manager UI
10 Full post-upgrade federation validation Both NSX GM + LM UI
Section 07

Key Gotchas and Architect Notes

  • 🟡 The old "LM must never lead GM" rule is outdated for 4.1.1+ and 4.2. Broadcom's official docs confirm that from NSX 4.1.1 onwards, and explicitly in 4.2, GM and LM can be upgraded in any order — federation sync is preserved across any version mix from 3.2 to 4.2. The N±1 rule only applied to NSX 4.1.0 and earlier. For VCF 5.x → 5.2, the prescribed sequence is still GM-first, but the reason is VCF BOM alignment and a resolved 4.1.x defect — not a hard architectural constraint. Always follow the official upgrade table for your exact VCF version: NSX 4.2 Federation Upgrade Guide.
  • 🟡 Edge nodes are managed under the LM domain in VCF. SDDC Manager handles NSX Edge node upgrades as part of the NSX component step. Do not manually upgrade Edge nodes via NSX UI — let SDDC Manager orchestrate it.
  • 🟡 GM config backup is your only recovery path. If the GM upgrade fails mid-way on a single-GM deployment, restoring from a pre-upgrade backup is the only supported recovery method. Verify backup integrity before starting.
  • 🔵 VCF BOM alignment is mandatory. SDDC Manager will only offer NSX versions that are part of its release BOM. If 4.2.3.1 isn't in the BOM of your target VCF release, SDDC Manager won't surface it — check the VCF release notes before planning your upgrade path.
  • 🔵 Cross-site config push is unavailable during GM upgrade. Plan your change window to account for the GM downtime period. Any configuration changes that need to propagate cross-site must be completed before or after — never during — the GM upgrade window.
  • 🟢 NSX 4.2.x improvements are worth the effort. The 4.2.x line brings significant improvements to federation replication reliability, VPC-mode support, and BGP graceful restart handling — all relevant for multi-site VCF deployments. The operational overhead of a careful upgrade sequence pays dividends in post-upgrade stability.
✓ Closing Note — Corrected

Federation upgrades reward preparation and accurate knowledge. The sequence — backup, validate federation health, upgrade GM manually, then let SDDC Manager handle LM — remains the right call for VCF 5.x deployments going to 4.2. But it's right because Broadcom's VCF upgrade table mandates it and there was a specific resolved defect in 4.1.x, not because "LM ahead of GM breaks federation." That old N±1 rule was retired in NSX 4.1.1.

Always verify the exact upgrade path for your version combination in the official Broadcom TechDocs Federation Upgrade Guide and cross-check with the VMware Interoperability Matrix before opening any change window.

Published on the VMware / Broadcom VCF Stack

VCF 5.x NSX 4.2 NSX Federation Global Manager Local Manager SDDC Manager Lifecycle Management vSAN

NSX Profiles: The Complete Reference from 4.x to VCF 9

NSX Deep Dive — Architecture Series

NSX Profiles: The Complete Reference
from 4.x to VCF 9

NSX VCF 9 Architecture Security April 2026 · 12 min read

If you have spent any time deploying NSX, you will have encountered profiles — and you will know that the variety and overlap between them can be genuinely confusing the first time you map them all out. Which profile controls my TEP VLAN? Where do I enforce anti-spoofing? Why do I have both a Segment Security profile and a SpoofGuard profile?

This post cuts through that confusion with a structured walkthrough of every profile type, what it does, and how it fits into the broader fabric of an NSX deployment. We also cover what has changed — and what has been confirmed as different — with the arrival of VCF 9 and NSX 9.

Think of NSX profiles as reusable configuration templates. Rather than configuring MTU, teaming, or security policy on each individual port or host, you define a profile once and apply it consistently across groups of objects. This is the foundation of a scalable, auditable NSX design.

1Fabric and infrastructure profiles

These profiles define how physical and virtual hardware form the NSX data plane. Getting these right is the foundation of any healthy deployment — everything from tunnel endpoint reachability to edge failover timing lives here.

Uplink profile

Fabric

Defines how a Transport Node connects to the physical network. Configures MTU, the transport VLAN for TEPs, and the teaming policy governing how physical NICs are used. Multiple named teaming policies can be defined within a single profile — essential for pinning TEP and workload traffic to different uplinks in a converged VDS deployment.

MTU · TEP VLAN · Failover / Load Balance Source

Transport Node Profile (TNP)

Fabric

A master template applied to an entire ESXi cluster. When attached, NSX automatically installs data plane components and configures the vDS on every host. A Sub-TNP variant handles stretched or multi-rack L3 topologies where different racks need distinct TEP VLANs.

Contains: Uplink Profile · Transport Zones · IP Assignment · vDS mappings

Edge Cluster Profile

Fabric

Controls liveness detection across Edge nodes. Configures BFD timers to determine how quickly the control plane declares an Edge node dead and triggers failover to its standby peer.

BFD probe interval · BFD declare dead multiple

Uplink profile tip: Named teaming policies are only usable with VLAN-backed segments. In a converged VDS environment, use them to pin TEP traffic to one NIC pair and workload traffic to another — without deploying a separate N-VDS. This is a common design pattern for management and compute clusters sharing a single VDS.

2Segment profiles

Segment profiles are applied to logical segments and their ports to govern how traffic is admitted, tracked, classified, and secured at the virtual port level. This is where the majority of day-to-day operational tuning happens.

IP Discovery profile

Segment

Determines how NSX learns the IP addresses of VMs attached to a segment. Uses ARP snooping, DHCP/DHCPv6 snooping, VMware Tools, and IPv6 ND snooping. Discovered bindings drive ARP suppression and are the authoritative source for DFW policy resolution.

Includes two trust models: Trust on First Use (TOFU), where the first discovered IP is pinned permanently, and Trust on Every Use (TOEU), where bindings age out dynamically.

Feeds: ARP suppression · SpoofGuard · DFW rule evaluation

MAC Discovery profile

Segment

Controls MAC address learning at the port level. MAC learning must be enabled for scenarios where more than one MAC appears behind a vNIC — nested hypervisors, Kubernetes CNIs, software bridges, and NFV appliances. Also controls whether a VM is permitted to change its own MAC address.

Critical for: Nested virt · Containers · NFV workloads

SpoofGuard profile

Segment

Actively enforces the IP and MAC bindings discovered by the IP Discovery profile. When a VM sends traffic whose source IP or MAC does not match its realised bindings, SpoofGuard drops the packet at the vNIC — before it reaches the segment. Enforced at both port and segment scope; both levels must pass.

Validates: MAC source · IP source · ARP/GARP/ND payload

Segment Security profile

Segment

Layer 2 security enforcement at the segment level. Controls BPDU filtering (preventing STP propagation into the overlay), DHCP server filtering (blocking rogue DHCP servers), and rate limiting on broadcast and multicast traffic. Works alongside SpoofGuard, not in place of it.

BPDU filter · DHCP server guard · Broadcast/multicast rate limits

QoS profile

Segment

Applies traffic marking and bandwidth shaping to tunnelled overlay traffic. Supports Layer 2 CoS (802.1p) and Layer 3 DSCP marking in trusted or untrusted mode. Bandwidth limiting sets average and peak (burst) transmit rates. Applies only to tunnelled traffic — intra-host VM-to-VM traffic bypasses this entirely.

Note: Does not apply to intra-host traffic

Common misconception: SpoofGuard is sometimes described as superseded by the Segment Security profile. This is inaccurate. They address different threat vectors — SpoofGuard validates IP and MAC identity at the packet level, while Segment Security controls Layer 2 protocol behaviour (BPDU, rogue DHCP, rate limits). In a hardened production environment both should be active. Broadcom TechDocs documents SpoofGuard as a first-class profile type in NSX 9.

3Security and gateway profiles

These profiles configure advanced inspection, threat prevention, and QoS behaviour for the gateway and firewall layers. They are consumed by the Distributed Firewall, NSX Gateway Firewall, and the Intrusion Detection and Prevention engine.

Context profile (L7)

Security

Enables the DFW to make decisions based on Layer 7 application identity rather than port and IP alone. NSX ships with a large library of FQDN-based, domain-based, and protocol-based application signatures — allowing policy that reflects business intent rather than network topology.

Used in: DFW rules as the App ID condition

IDS/IPS profile

Security

Configures the NSX Distributed IDS/IPS engine. Defines which signature sets are active, severity thresholds, and whether the engine runs in detect-only or active-block mode. Signatures can be scoped by CVE severity, CVSS score, or affected product family.

Modes: Detect only · Detect and Prevent

Malware Prevention profile

Security

Controls how the NSX malware detection engine extracts files from east-west flows and submits them for sandbox analysis via NSX Advanced Threat Prevention. Configures file type filters and the action taken on a positive verdict.

Requires: NSX Advanced Threat Prevention licence

Gateway QoS profile

Security

Applies bandwidth shaping to north-south traffic flowing through Tier-0 or Tier-1 gateway uplinks. Distinct from Segment QoS which operates on east-west overlay traffic. Used to enforce CIR and burst rates per gateway uplink — important in multi-tenant environments to prevent a single tenant saturating shared edge uplinks.

Scope: Gateway uplink interfaces only


4What changed: NSX 4.x vs VCF 9

The profile types themselves are largely consistent across versions. What VCF 9 changes is the ownership and governance model around them. The single most important confirmed change: NSX is no longer available as a standalone product.

AreaNSX 4.xVCF 9 / NSX 9Status
Deployment model Deployable standalone or within VCF Exclusively delivered within VCF stack; standalone deployment blocked Breaking change
Lifecycle management NSX upgradeable independently via NSX Manager UI All upgrades orchestrated by SDDC Manager as part of VCF lifecycle Removed
Policy API Policy API preferred; Manager API legacy objects still accessible Policy API is the exclusive standard; deprecated and removed APIs formally catalogued in NSX API Guide Tightened
Transport Node Profile Managed via NSX Manager UI or SDDC Manager TNP lifecycle owned by SDDC Manager; out-of-band changes risk drift flagged by the platform VCF-owned
Edge deployment Edge VMs deployed via SDDC Manager JSON/UI New vCenter UI wizard available; Tier-1 replaced by Transit Gateway in VPC networking model New capability
Networking models Segment networking only (T0 / T1 / Segments) Two models: Segment Networking (unchanged) and VPC Networking (new cloud-native model) VPC model added
IWA authentication Integrated Windows Authentication supported IWA deprecated; LDAP/S is the replacement Deprecated
Segment profiles Full profile set (IP Discovery, MAC, SpoofGuard, Security, QoS) All segment profile types confirmed present and unchanged in NSX 9 Confirmed
Certificate management Manual certificate rotation Upgrade precheck for expiring Transport Node SSL certs (90-day warning); CARR script integration New precheck

Architect's note on the VPC networking model: VCF 9 introduces VPC Networking alongside the familiar Segment Networking model. Segment Networking remains fully supported and is the right choice where centralised admin control is required. VPC Networking aligns with public-cloud consumption patterns and enables self-service for tenants via NSX Projects. The Transit Gateway is the key new construct — interconnecting VPCs and connecting them to physical infrastructure either via a Tier-0 (centralised, requires Edge VMs) or directly via an external VLAN (distributed connectivity, no Edge VMs required).


5Practical design principles

Keep Uplink Profiles minimal

Create one Uplink Profile per physical hardware pattern. If all hosts share the same NIC speed, TEP VLAN, and MTU, a single profile should cover your entire compute fabric. Proliferating profiles for minor differences makes lifecycle management harder under SDDC Manager.

Treat TNPs as immutable once applied

In VCF 9, SDDC Manager owns the TNP lifecycle. Make changes through a change-controlled TNP update process via SDDC Manager — not ad-hoc in the NSX UI. Configuration drift between the TNP and host state is now flagged by the platform and will block upgrade operations.

Always pair IP Discovery with SpoofGuard

IP Discovery alone is passive — it discovers and logs. SpoofGuard is what enforces. For production segments carrying sensitive workloads, enable both. Use TOFU mode for stable workloads; use TOEU for dynamic environments where IP churn is expected and binding permanence would cause operational friction.

Enable MAC learning intentionally

MAC learning is off by default for good reason — enabling it on a general-purpose workload segment weakens address-based policy enforcement. Enable it only on segments explicitly designed for nested hypervisors, containers, or network appliances. When you do, enable SpoofGuard in parallel to compensate.

Summary

NSX profiles have remained structurally consistent from NSX-T 3.x through NSX 4.x and into VCF 9. The fabric layer (Uplink, TNP, Edge Cluster), segment layer (IP Discovery, MAC, SpoofGuard, Segment Security, QoS), and security layer (Context, IDS/IPS, Malware Prevention, Gateway QoS) are all present in the current release. What VCF 9 changes is not the profile taxonomy — it is the ownership model. SDDC Manager is now the authoritative orchestrator for profile lifecycle, the Policy API is the only sanctioned management path, and NSX can no longer be deployed or upgraded in isolation. For architects migrating from 4.x, the profile knowledge transfers directly; the operational shift is in governance, not configuration.

Breaking the Broadcom Shackles: Is Escaping NSX, Avi & vSAN Actually Viable?

VMware · Broadcom · Cloud Infrastructure

Breaking the Broadcom Shackles: Is Escaping NSX, Avi & vSAN Actually Viable?

When you've built your entire operational DNA around VMware's holy trinity, leaving isn't a migration — it's an architectural divorce.

Let's be honest with ourselves. The question everyone is asking post-Broadcom acquisition isn't "should we leave VMware?" — it's "can we actually leave, given how deep we are?" If you're running the full holy trinity of NSX, Avi Networks, and vSAN, you're not just using software. You've woven Broadcom's proprietary logic into every layer of your stack.

The frustration is completely understandable. Broadcom's new per-core VCF pricing model has landed like a grenade in most infrastructure budgets. But pricing anger and actual migration viability are two very different conversations — and conflating them leads organisations to make rash decisions they'll regret for years.

So let's have the honest conversation. Component by component.

The stickiest trap: NSX & Avi Networks

If you had to rank the VMware components by "escape difficulty," NSX belongs at the top. It's not just a network overlay — it's where your entire security posture lives. Micro-segmentation rules, distributed firewall policies, VXLAN/Geneve encapsulation logic, and your east-west traffic inspection are all baked into NSX constructs that exist nowhere else in the same form.

The core problem When you move to KVM, Nutanix AHV, or any non-VMware hypervisor, your NSX security groups, firewall rules, and Avi load-balancing configuration don't "export." There is no migration wizard. You are re-architecting your entire network security posture from a blank canvas.

For Avi specifically, most organisations looking to exit are reverting to F5 BIG-IP or NGINX as a replacement. Both are credible. Neither gives you the tight "one-pane-of-glass" integration that VCD provided with Avi sitting natively within the platform. You are trading operational simplicity for vendor independence — and that trade-off has a real cost in engineer hours.

The most technically interesting alternative gaining traction in the Kubernetes-forward space is Cilium with eBPF for networking and security policy enforcement. It's genuinely powerful, but it requires a meaningful architectural pivot toward container-native thinking. If your workloads are still primarily VM-based — and in most enterprise environments they are — Cilium isn't your NSX replacement. Not yet.

The storage wall: vSAN is a data migration project in disguise

People often underestimate vSAN migration because it looks like a software problem. It isn't. vSAN is software-defined storage running on your local server disks, which means your data is physically distributed across your ESXi hosts in a proprietary format. To leave, you need somewhere to put the data first.

Option A

Physical SAN as swing space

Buy NetApp or Pure Storage as a temporary migration target. Capital-intensive and requires major procurement lead time.

Option B

Host-by-host HCI migration

Migrate workloads incrementally to Nutanix AOS. Operationally gruelling and typically spans 12–18 months minimum.

Option C

Stay put, reduce tier

Retain vSAN but negotiate down to the minimum required licensing tier. Reduce costs without the migration risk.

The brutal reality is that most organisations aren't doing a full vSAN exit. They're staying put because the capital cost of a migration — both hardware and engineering time — frequently exceeds the increased Broadcom licensing cost for the first three years. The maths doesn't work for a full rip-and-replace until you're operating at meaningful scale.

The platform parity problem

This is the dimension that gets the least attention in migration conversations, and it's arguably the most important one for service providers running VCD environments.

VCD is enterprise-polished in a way that most alternatives simply aren't. Tenant self-service portals, billing integration, service catalog management, role-based access across multiple organisations — VCD handles all of this in a way that feels coherent because it was designed as a complete platform. The sum is greater than the parts.

The gap no one talks about Apache CloudStack is technically robust and genuinely capable, but operating it at enterprise scale requires significantly more engineering investment to maintain feature parity. You stop paying Broadcom in dollars, and you start paying your engineering team in overtime. The budget line changes; the total cost of ownership doesn't necessarily.

OpenStack is a similar story — powerful, flexible, cloud-native by design, but it carries substantial operational overhead. Proxmox is excellent for smaller environments but simply isn't built for the multi-tenancy and scale that VCD handles natively.

"Broadcom knows exactly how hard it is to leave NSX and vSAN. That stickiness is precisely why they feel confident in the current pricing model."

The middle ground: strategic shrinking

Most organisations that are handling this well aren't planning a 100% VMware exit. They're pursuing what I'd call strategic shrinking — a deliberate, phased approach to reducing VMware's footprint and licensing exposure without the risk of a full platform migration.

  1. Freeze the VMware footprint Stop growing the VCF/VCD environment. No new compute added under Broadcom licensing. Existing capacity becomes legacy.
  2. Route new workloads elsewhere Any new customer deployment, new internal project, or greenfield service goes onto Nutanix or a KVM-based platform from day one.
  3. Negotiate the legacy island down Keep the NSX/vSAN stack for applications that are genuinely too complex to migrate, but work aggressively on getting your licensing tier to the minimum viable configuration.
  4. Build migration competency gradually Use lower-stakes workload migrations to build your team's capability on the target platform before tackling the genuinely difficult stuff.

The verdict: it depends entirely on your scale

Organisation type Full exit viable? Why
Small / mid-tier MSP
Under ~500 cores
Unlikely Migration cost (hardware + engineering time) typically exceeds the Broadcom tax for the first 3 years. The maths don't work.
Mid-market enterprise
500–2,000 cores
Strategic shrink Partial exit viable. Freeze and redirect new workloads. Negotiate the legacy core down.
Large enterprise / global provider
5,000+ cores
Yes, long-term At this scale, the Broadcom tax is high enough that hiring a dedicated OpenStack or Nutanix engineering team becomes the cheaper long-term play.

What should you actually do right now?

Before making any platform decisions, do the one thing most organisations skip: build a proper total cost of ownership model. Not just the licensing delta — the full picture. Engineering time to rebuild NSX equivalent security posture. Data migration hardware and project cost. Training and upskilling on the target platform. Operational tooling gaps you'll need to fill.

When you do that analysis honestly, the decision usually becomes clearer. For smaller providers, the answer is often "negotiate hard with Broadcom and reduce your licensing tier aggressively." For larger ones, it's "start the journey now, because the longer you wait, the more the dependency deepens."

Either way, the worst strategy is paralysis. The organisations that will be in the best position in three years are the ones that made a deliberate choice — exit, shrink, or stay — with eyes open, not the ones still waiting to see how the pricing model evolves.

Broadcom is betting you'll wait. Don't prove them right.

Are you currently evaluating a VMware exit strategy?

I'd be interested to hear where you are in the process — whether you're still on a legacy perpetual/VCPP model, or you've been quoted on the new per-core VCF pricing. Drop your situation in the comments below.

NSX T1 Gateway — SR vs DR and Edge Cluster Residency

NSX-T T1 Gateway: Does a T1 Without an SR Still Live on the Edge Cluster?

📅 Published by Farrukh Hanif 🏷 NSX-T  |  VMware  |  VCF  |  Networking
If a T1 Gateway has no Services Router (SR), where does north-south traffic actually go? And does the T1 even touch the Edge cluster? This is one of those NSX-T concepts that catches people out — and understanding it is fundamental to designing efficient overlay networks.

The Two Components of Every NSX-T Gateway

Every Tier-0 and Tier-1 Gateway in NSX-T is logically made up of two distinct functional components:

  • Distributed Router (DR) — runs on every transport node (hypervisor) that hosts a connected workload segment. It handles east-west routing entirely in the data plane of the host, with no hairpinning through a centralised appliance.
  • Services Router (SR) — a centralised component instantiated on the Edge cluster. It handles stateful services that cannot be distributed: NAT, load balancing, stateful gateway firewall, and VPN termination.

The key insight is this: the SR is optional on a T1. NSX-T will only instantiate a T1 SR on the Edge cluster if you configure a service that actually requires it. Without such a service, the T1 remains a purely distributed router.

Architecture Diagram

Edge cluster Physical network / underlay T0 SR Edge cluster — physical uplink T1 SR NAT / LB / FW Hypervisor / transport nodes (distributed plane) T0 DR Distributed on all hypervisors — local handoff point for N/S traffic T1 with SR (stateful services) T1 DR Distributed — E/W routing on hypervisor to T1 SR VM-A VM-B N/S: VM → T1 DR → T1 SR (Edge) → T0 SR (Edge) → physical T1 without SR (no stateful services) T1 DR only Distributed — no Edge presence whatsoever local handoff VM-C VM-D N/S: VM → T1 DR → T0 DR (same hypervisor — no Edge hop for T1) → T0 SR (Edge) → physical Key Takeaway The T0 DR runs distributed on every hypervisor alongside the T1 DR. Without a T1 SR, N/S traffic hands off T1 DR → T0 DR locally on the same host — no Edge hop for T1. The T0 SR on the Edge cluster provides the physical uplink. East-west between T1 segments stays on the hypervisor entirely in both cases.

The Short Answer

No — a T1 without an SR does not reside on the Edge cluster at all. The T1 DR continues to operate on every hypervisor hosting a connected segment, but there is no Edge component, no Edge Edge Node allocated, and no traffic path through the Edge for that T1.

So How Does North-South Traffic Flow?

This is where the design becomes elegant. When a VM on a T1-without-SR needs to reach an external destination, the forwarding path is entirely local to the hypervisor until it hits the T0's Edge SR:

T1 with SR (stateful services enabled)

VM → T1 DR (local hypervisor) → T1 SR (Edge cluster — NAT/LB/FW processed here) → T0 SR (Edge cluster) → Physical network

T1 without SR (no stateful services)

VM → T1 DR (local hypervisor) → T0 DR (same hypervisor — no Edge hop for T1) → T0 SR (Edge cluster — uplink to physical) → Physical network

The packet never leaves the hypervisor for T1 routing. The T0 DR, also running distributed on every transport node, receives the handoff from the T1 DR locally and then forwards to the T0 SR on the Edge for the external uplink. This is a significant efficiency gain — it means one fewer hop through the Edge cluster for every north-south flow from that T1.

When Does a T1 Get an SR?

NSX-T automatically instantiates a T1 SR on the Edge cluster when you configure any of the following on that T1:

Service SR Required? Reason
NAT (SNAT/DNAT) Yes Stateful connection tracking required
Load Balancing Yes Connection state must be centralised
Gateway Firewall (stateful) Yes Session table cannot be distributed
IPsec VPN termination Yes Crypto and tunnel state is centralised
Simple routing (no services) No DR handles all east-west; T0 handles N/S uplink
Distributed Firewall (DFW) No DFW is kernel-level on each hypervisor, independent of gateways

East-West Traffic: Always Fully Distributed

Regardless of whether an SR exists or not, east-west routing between segments connected to the same T1 is always handled entirely within the T1 DR on the local hypervisor. If VM-A on Segment-A pings VM-B on Segment-B, and both segments are on the same T1, the packet goes:

VM-A (Hypervisor-1) → T1 DR (Hypervisor-1, routes to Segment-B) → Geneve tunnel (if VM-B is on a different host) → VM-B

No Edge cluster involvement at all. This is the whole point of the distributed routing architecture — it keeps high-volume lateral traffic off the Edge nodes entirely.

💡 Design Principle The SR is not a "better" or "more capable" version of the DR. It is a separate, narrowly-scoped component that exists only to host stateful services that fundamentally cannot be distributed. Avoid treating it as the default routing path.

Practical Design Guidance

Do not add an SR to a T1 unless you need it

Over-provisioning SRs is a common anti-pattern. Every T1 SR you create consumes Edge Node capacity — memory, CPU, and a pinned Edge Node path. If your workloads just need routing between segments and outbound internet access (handled by NAT on the T0), there is no reason to instantiate a T1 SR.

NAT design: T0 or T1?

You can configure NAT on either the T0 or a T1. If you put NAT on the T0, your T1s can remain SR-free. This is often the right call in smaller environments. In larger multi-tenant designs, T1-level NAT gives you isolation per-tenant at the cost of each T1 needing its own SR on the Edge.

Edge cluster sizing implications

Each T1 SR is placed on an Edge Node in the Edge cluster. If you have 20 T1 gateways all with SRs, your Edge cluster needs to be sized to handle the aggregate stateful service load from all 20. This is a critical input to Edge cluster sizing — count your SRs, not just your T1 count.

⚠ VCF Lab Note In a VCF 9 home lab with constrained Edge cluster resources (such as deploying Edge VMs on an R630 with 368GB RAM), keeping T1 gateways SR-free wherever possible will significantly reduce your Edge node memory footprint and leave headroom for additional NSX services.

Verifying SR Status via NSX Manager

You can quickly confirm whether a T1 has an SR instantiated from the NSX Manager UI:

  1. Go to Networking > Tier-1 Gateways
  2. Click the three-dot menu on your T1 and select View and manage gateway
  3. Under Routing, check whether an SR is listed and which Edge Node it is placed on

Via the NSX CLI on an Edge Node, you can also run:

# On the Edge Node NSX CLI
get logical-routers

# Look for your T1's SR component (VRF type: SERVICE_ROUTER_TIER1)
# If only a DISTRIBUTED_ROUTER_TIER1 appears for your T1, no SR is instantiated

Summary

The presence or absence of an SR on a T1 is not about capability — the T1 DR always exists and always routes. The SR is strictly a function of whether you need stateful services. Without stateful services, the T1 is a lean, fully distributed router with no Edge footprint. North-south traffic traverses the T1 DR locally and is handed off directly to the T0 DR on the same hypervisor, which then forwards to the T0 SR on the Edge for the physical uplink.

Understanding this architecture is essential for anyone designing NSX-T environments — whether you are sitting the VCIX-NV exam, designing a multi-tenant VCF deployment, or simply trying to understand why your Edge nodes are under pressure.

💬 Questions or corrections? Drop a comment below or reach out on LinkedIn. I document real lab deployments, architecture decisions, and exam prep as part of my journey to a Principal Cloud Architect role in 2026.
FH
Farrukh Hanif

NSX/VCF Engineer at NatWest Group with 15+ years of VMware, NSX-T, VCF, and cloud infrastructure experience. VCIX6-NV | VCP-VCF9 | VCAP-NV Design | CKA | CKS | AWS SA Pro. Building a VCF 9 home lab and documenting the journey. Connect on LinkedIn.

VCF 9 Home Lab | Embedded vIDM (viDB) --- AD Integration, Users, Groups & NSX SSO

VCF 9 Home Lab | Embedded vIDM (viDB) — AD Integration, Users, Groups & NSX SSO 📅 May 2026  |  🏷️ VCF 9 Home Lab Series  ...