Logo for NVIDIA

Senior Solutions Architect - AI Factory Deployment

Role overview

Qualifications

  • Bachelor's degree or equivalent in Computer Science, Mathematics, Engineering, Physics, or related field.
  • 6+ years of experience managing Linux-based HPC/distributed systems or AI/ML environments, with hands-on workloads on multi-GPU/multi-node clusters.
  • Deep knowledge of collective communication patterns (AllReduce and AllToAll) and practical experience with NCCL in ML/LLM training workflows (PyTorch or TensorFlow).
  • Proficiency in Python and Shell/Bash scripting, plus experience with benchmarking and observability tooling (metrics, logs, dashboards) to automate and monitor performance.

Responsibilities

  • Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters, ensuring NCCL/collectives and distributed training framework configurations.
  • Own execution of key AI/LLM benchmarks: setup, orchestration, result collection, analysis, and issue resolution for underperforming or failed jobs.
  • Build and improve observability for AI factories, including metrics, logs, traces, and dashboards; develop automation for benchmarking and regression checks.
  • Collaborate across hardware, software, networking, data center, and product teams to prepare AI factories for customer use; contribute to documentation and readiness collateral.

About the company

NVIDIA logo

NVIDIA

Semiconductors

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Company details

Company typeXLarge
IndustrySemiconductors
Company size10001

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

We are seeking an ambitious Senior Solutions Architect - AI Factory Deployment to join our NVIDIA Infrastructure Specialists team in Santa Clara! This role is uniquely positioned to develop, deploy, and validate AI factories end to end. You will focus on running and debugging AI/LLM workloads and benchmarks on Linux-based GPU clusters, using NCCL and collectives like AllReduce and AllToAll to improve performance and scalability.

As part of our world-class team, you will bring to bear observability and automation to improve benchmarks and validation. You will serve as the expert when workloads or benchmarks do not perform flawlessly. You will collaborate across NVIDIA to ensure AI factories are prepared for customers, validating hardware and software for modern AI deployments.

What You Will be Doing:

  • Set up, adjust, and verify AI factory environments across multi-GPU and multi-node Linux clusters.

  • Ensure configurations align with guidelines for NCCL, collectives, and distributed training frameworks.

  • Own the execution of key AI/LLM benchmarks, including setup, orchestration, result collection, and analysis.

  • Investigate and resolve issues when training jobs or benchmarks fail, hang, or underperform.

  • Build and improve observability for AI factories (metrics, logs, traces, dashboards) to understand workload behavior and system health.

  • Develop automation (Python, Shell) for running benchmarks, collecting results, and performing regression checks

  • Examine communication patterns and NCCL usage for AI/LLM workloads, concentrating on collectives such as AllReduce and AllToAll.

  • Recommend changes to job configuration, parallelism strategies, and cluster settings to improve throughput, latency, and scaling efficiency.

  • Work closely with hardware, software, networking, datacenter, and product teams to prepare AI factories for customer use.

  • Contribute to documentation, guidelines, and readiness collateral that support internal collaborators and customer-facing teams.

What We Need to See:

  • Bachelor’s degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field.

  • More than 6+ years of experience managing Linux-based systems in HPC, distributed systems, or extensive AI/ML settings.

  • Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters, with practical knowledge of NCCL.

  • Solid grasp of collective communication patterns, particularly AllReduce and AllToAll, and how they are applied in contemporary ML/LLM training.

  • Familiarity with LLM training and/or inference workflows using frameworks such as PyTorch or TensorFlow.

  • Proficiency with Python and Shell/Bash for scripting, automation, and tooling.

  • Experience with benchmarking (crafting, executing, and interpreting performance benchmarks).

  • Comfortable working with observability data (metrics, logs, dashboards) to troubleshoot and optimize complex distributed workloads.

  • Strong communication skills and the ability to work effectively with cross-functional teams.

Ways to Stand Out From the Crowd:

  • Experience with AI factory or large-scale AI infrastructure build, deployment, or operations.

  • Background in HPC performance engineering, SRE, or systems performance analysis for GPU-accelerated environments.

  • Familiarity with observability stacks (e.g., metrics/monitoring, logging, tracing systems) used for large distributed systems.

  • Experience building automation and CI-style pipelines for running and validating benchmarks at scale.

  • Demonstrated desire to use AI to solve practical problems, improve workflows, and guide data-driven decisions.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until May 3, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
Β·

Solutions Architect Related jobs

Other jobs at NVIDIA

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.