NVIDIA
Principal System Power Management and Performance Architect
US, CA, Santa Clara · Posted 1h ago
Job Description
NVIDIA builds the silicon behind AI, accelerated computing, and graphics. Every watt of performance and every degree of thermal headroom traces back to decisions made in power, performance, and thermal architecture. We are the Silicon Co-Design Group (SCG). We identify, own, and drive system-level co-design ideas. We start with initial concepts and advance to product differentiation across NVIDIA's roadmap.
We are hiring a Principal System Power Management and Performance Architect who operates at the ambiguous boundary where workload behavior, silicon capabilities, firmware policies, and platform constraints collide, and who turns that ambiguity into architecture that survives across multiple silicon generations. SCG scope spans architecture, design, software, operations, platforms, and productization. This role shapes system, platform, and data center features and behavior, and partners with teams across NVIDIA.
What You'll Be Doing:
The work here is rarely well-defined when it arrives. You will be given problems that appear to be performance gaps or power anomalies and encouraged to build a framework for solving them, not just tackle a single instance.
Define the multi-generation roadmap for system-level power and performance features, grounded in prototyping, use-case analysis, and cost/benefit trade-offs across segments. You will decide what problems are worth solving and why.
Own the architecture and integration strategy for HSIO power management, DVFS, P-states, and low-power features. Your decisions improve product performance, power, and reliability across product lines — not just the current program.
Lead system-level boot and IST architecture defining how power and clock domains initialize, sequence, and recover across complex multi-IP systems where the interaction space is large and the failure modes matter.
Drive power management strategy at datacenter scale, including rack-level power telemetry and platform state coordination in high-performance environments.
Identify and redesign system-level processes that break under new product requirements—boot/reset flows, low-power entry/exit sequences, and control-system policies across IPs. When a prior design no longer holds, you diagnose why and architect what comes next.
Serve as the multi-functional technical authority across architecture, ASIC, board/platform, and firmware teams. You improve the power-performance trade-offs by influencing decisions made by teams you do not control.
What We Need to See:
We are calibrating for candidates who can carry ambiguous product-level power and performance problems from concept through silicon correlation to release trade-off, and who can show the work!
BS or MS in EE/CE, or equivalent experience proven through the work itself
15+ years of system architecture, development, and validation experience with a strong focus on power and performance-per-watt optimization in datacenter or high-performance platforms.
A verifiable history of guiding architecture decisions that shipped across multiple silicon programs. Be ready to walk through a case where you challenged a roadmap direction, explain your reasoning from first principles, and describe what changed.
Extensive knowledge of HSIO power management, DVFS, P-states, system boot flows, and low-power entry/exit architecture, including how they interact with IP, firmware, and platform layers. The expectation is not pattern-matching to prior solutions; it is knowing why the design space is built the way it is.
Strong fundamentals in low-power design, power management techniques, and rack-level performance architecture.
Ways to stand out from the crowd:
We are not looking for breadth of exposure. We are looking for depth of ownership in areas where the design space is hard, and the artifacts prove it!
Practical, hands-on depth in at least two of the following: Preference is given to candidates who have debugged these features in silicon and then architected the next-generation solution in any of the two areas: system power management, robust boot flow optimizations, HSIO, and memory power management.
A power management architecture or methodology you introduced that survived multiple silicon generations and was adopted beyond your immediate team. We want to understand the original problem, why prior approaches failed, how you structured the solution, and what the adoption path looked like.
A demonstrated AI/agentic workflow practice: Not just tool use, but a disciplined approach to compressing engineering velocity with AI while applying strong judgment to validate and refine results. Be prepared to describe where you've found AI dangerous, not just useful.
Visible technical leadership through patents, publications, conference presentations, standards participation, or broader industry recognition for work that influenced products, roadmaps, or engineering practice.
Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/
#LI-Hybrid
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 232,000 USD - 368,000 USD.You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.