Skip to content
NVIDIA 5.8 trillion won investment in optics: implementation strategy to change AI data center bottlenecks
← Back to blog

NVIDIA 5.8 trillion won investment in optics: implementation strategy to change AI data center bottlenecks

Development·7 min read

Based on NVIDIA's signal of large-scale investment in Lumentum and Coherent, the infrastructure team organized a practical framework to inspect and mitigate optical component risks within 90 days.

NVIDIA 5.8 trillion won optical investment: implementation strategy to change AI data center bottlenecks

Publication date: 2026-03-03 | Category: Development information

1) Problem definition

Target readers are CTOs, platform teams, and infrastructure purchasing personnel who operate or procure AI infrastructure. The problem is that the supply risk of optical network components (transceivers, lasers, optical interconnects) that actually use the GPU becomes a bottleneck faster than the performance of the GPU itself. NVIDIA's announcement of a total investment of $4 billion (approximately 5.8 trillion won) in Lumentum and Coherent is interpreted as a strategic action to preemptively close this bottleneck. This article is not a simple news summary, but presents standards for supply chain and architecture decisions that can be implemented within 90 days. However, internal contract terms/private price negotiations of individual companies are excluded from the scope.

2) Evidence and comparison

The essence of this issue is that the center of gravity has shifted from “competition to secure GPU” to “competition to secure optical capacity.” Even within the same AI cluster, the effective throughput varies greatly depending on the network design and supply contract structure.

AccessAdvantagesLimitRecommendation status
GPU-centric procurementSimple decision-making, fast initial orderingRack unit idle occurs when optical components are delayedPoC/Short-term benchmark
Optical+GPU simultaneous capacity contractStabilization of actual operation rateInitial negotiations are complicated, supplier diversification is requiredCommercial services/Large-scale learning
Multi-vendor optical standardizationMitigating specific vendor risksIncreased compatibility verification/operation difficultyOrganization operating in 2 or more regions
  • Cost: Delay cost (idle power/opportunity cost) is higher than GPU unit price. There are many sections that increase the total cost.
  • Time: Optical component lead times are now more likely to dictate deployment schedules.
  • Accuracy/Performance: Eliminating cluster network bottlenecks before improving model quality is directly reflected in perceived performance.
  • Difficulty: The cost of cross-organizational alignment is high as procurement, network, SRE, and finance must move together.

3) Step-by-step execution method

  1. D+1~7: Bottleneck visualization — Quantifies the optical module/switch port margin ratio compared to the number of GPUs per cluster, and sets lead time risk to High/Medium/Low. Classify.
  2. D+8~21: Create two or more procurement scenarios — Compare the total cost of ownership (TCO) and delay risk of (A) single vendor lock-in and (B) multi-vendor mix.
  3. D+22~45: Verification of architectural suitability — Check the bottleneck section of the optical layer through a load test according to the learning/inference traffic pattern.
  4. D+46~70: Enhanced contract conditions — In case of supply delay, alternative supply/penalty/priority allocation provisions are reflected in the contract.
  5. D+71~90: Operational transition gate — If uptime, delay, and failover time criteria are not exceeded, staged rollout is halted and revalidated.
#capacity gate example
if optical_buffer_weeks < 6 or network_p95_latency_ms > target:
  block_cluster_scaleout()
  activate_backup_vendor_plan()

4) Mistakes/Pitfalls

  1. Trap: Assuming that it is over once GPU supply is secured
    Prevention: In the quarterly plan Required inclusion of optical component lead time KPI
    Recovery: Minimize idle section by rearranging rack expansion priorities
  2. Pitfall: Loss of negotiating power due to dependence on a single optical vendor
    Prevention: Constant certification/compatibility testing by at least 2 vendors Maintain
    Recover: Immediately convert high-risk SKUs to alternative specifications
  3. Pitfall: Misdiagnosis of network bottlenecks as model/software issues
    Prevention: Learning step time and network metrics together Monitoring
    Recovery: Replace/relocate the optical link in the bottleneck section first

5) Execution Checklist

  • GPU ordering plan and optical component ordering plan are tied to the same calendar
  • Weekly check of optical component lead time (weekly) and inventory buffer (weekly)
  • Documented alternative procurement path assuming single vendor failure
  • View network p95 delay, packet drop, and cluster uptime in one dashboard
  • The finance, infrastructure, and service teams jointly approve the expansion Go/No-Go criteria

Definition of Done: If “0 cluster expansion delays due to optical component delays + target operation rate achieved + at least 1 rehearsal of alternative procurement plan completed” for two consecutive quarters Completed.

6) Reference

7) Author Viewpoint

I do not view this news as “another investment article.” The key to the AI ​​infrastructure race in 2026 is not GPU quantity, but optical supply stability, which translates GPUs into actual throughput. Therefore, the infrastructure team must immediately switch from ‘GPU-centric KPI’ to ‘cluster effective utilization KPI’. Conversely, organizations with small traffic and slow expansion speed do not need to go to a complex multi-vendor system right now. However, it is no longer an option to at least stipulate alternative scenarios in both contract and operation in case of supply disruption.

Share this article

Related articles

Take the AQ test

See your AI capability in three minutes. Assess recognition, utilization, verification, integration, and ethics at once, then receive practical insights.

Start the free AQ test