Senior Solutions Architect – Large Scale Neural Networks Inference — NVIDIA (ufficio Zurich)
- Location
- Zürich
- Contract
- full-time
- Posted
- 3 days ago
Role overview
We are seeking a Senior Solutions Architect with deep expertise in large-scale neural network inference and a proven ability to lead technical collaboration with frontier AI labs and enterprises deploying AI at scale across EMEA.
In this role, you will define the technical direction for AI inference across EMEA by identifying critical bottlenecks and driving the development of scalable, high-impact solutions.
- We are seeking a Senior Solutions Architect with deep expertise in large-scale neural network inference and a proven ability to lead technical collaboration with frontier AI labs and enterprises deploying AI at scale across EMEA.
- In this role, you will define the technical direction for AI inference across EMEA by identifying critical bottlenecks and driving the development of scalable, high-impact solutions.
Company and context
- By aligning key team members within NVIDIA and customer organizations, you will influence strategic technology decisions to develop the deployment of next-generation AI inference at scale. What you will be doing:
- Lead the inference strategy for a portfolio of EMEA AI Natives customers, guiding engagements from initial proof of concept to production-scale deployments.
- Identify inference challenges across customer deployments including latency, efficiency, cost per token, memory utilization, and low-latency networking.
- Architect and optimize high-performance inference pipelines using NVIDIA Dynamo, TensorRT-LLM, vLLM, SGLang, and other inference backends, improving GPU utilization and AI cluster efficiency.
- Translate customer insights and deployment patterns into actionable product feedback that develops the roadmap for NVIDIA stack such as Dynamo, TensorRT-LLM, and NIM. What we need to see:
- MS or PhD in Computer Science, Engineering, or equivalent experience in the field.
- 8+ years in AI/ML infrastructure, with deep expertise in LLM/VLM inference optimization and production deployment at scale.
- Deep understanding of transformer inference acceleration: quantization (INT4/FP8), speculative decoding, disaggregated inference, continuous batching, KV cache optimization, and WideEP for MoE models.
- Understanding of GPU memory hierarchies and low-latency networking along with their influence on inference performance.
- Proven track record to lead technical initiatives.
Additional details
- By aligning key team members within NVIDIA and customer organizations, you will influence strategic technology decisions to develop the deployment of next-generation AI inference at scale. What you will be doing:
- Translate customer insights and deployment patterns into actionable product feedback that develops the roadmap for NVIDIA stack such as Dynamo, TensorRT-LLM, and NIM. What we need to see:
- Excellent communication skills, effective with research scientists, infrastructure engineers, and executive team members. Ways to stand out from the crowd:
Notes and original content
- By aligning key team members within NVIDIA and customer organizations, you will influence strategic technology decisions to develop the deployment of next-generation AI inference at scale.
- What you will be doing:
- Translate customer insights and deployment patterns into actionable product feedback that develops the roadmap for NVIDIA stack such as Dynamo, TensorRT-LLM, and NIM.
- What we need to see:
- Excellent communication skills, effective with research scientists, infrastructure engineers, and executive team members.
- Ways to stand out from the crowd:
Questions about this listing
What salary does NVIDIA (ufficio Zurich) offer for this role?
NVIDIA (ufficio Zurich) lists CHF 113'500 - 172'000 gross per year for this position in Zürich. This is the salary published in the original listing (or, when the employer omits a figure, a realistic estimate for the role and sector) — the calculator on this site converts it to your actual net take-home once cross-border tax and social contributions are applied.
Is this a full-time role, and what type of contract does NVIDIA (ufficio Zurich) offer?
This listing is a full-time position. The contract type shown here comes directly from the employer's original posting; always confirm exact hours, notice period and probation length with NVIDIA (ufficio Zurich) during the application process, since these details can vary by role even within the same contract category.
Do I need a cross-border work permit for a role in Zürich?
EU/EFTA residents living in the border zone of the country adjoining Canton Zürich can apply for a G permit; the Swiss employer files it with that canton's migration office after the contract is signed. Border-zone rules and processing times vary by neighbouring country and canton, so confirm the specifics with Zürich's cantonal migration office or with HR during the application.
How do I apply for this position at NVIDIA (ufficio Zurich)?
Use the "Apply now" button on this page — it links directly to NVIDIA (ufficio Zurich)'s original listing at nvidia.wd5.myworkdayjobs.com, so your application goes straight to the employer's own applicant-tracking system. Frontaliere Ticino does not collect or forward applications itself.