Back to jobs

Senior Director, Data Science Manager

New York, NY, United States

Pay
Salary not listed in the saved posting
Work setup
Unconfirmed
Employment
Unconfirmed
Apply at BNY Mellon

What you’ll bring

All qualifications

Preferred experience

  • Experience within financial services or other highly regulated industries.
  • Expertise supporting AI platforms subject to stringent security, compliance, audit, and operational resilience requirements.
  • Experience with large-scale GPU environments, DGX platforms, SuperPOD architectures, InfiniBand, high-performance storage, or enterprise AI infrastructure.
  • Familiarity with modern AI serving and orchestration technologies such as Triton, vLLM, NVIDIA NIM, KServe, Seldon, LangGraph, vector databases, graph databases, and multi-agent frameworks.
Qualification wording
Experience within financial services or other highly regulated industries.
Expertise supporting AI platforms subject to stringent security, compliance, audit, and operational resilience requirements.
Experience with large-scale GPU environments, DGX platforms, SuperPOD architectures, InfiniBand, high-performance storage, or enterprise AI infrastructure.
Familiarity with modern AI serving and orchestration technologies such as Triton, vLLM, NVIDIA NIM, KServe, Seldon, LangGraph, vector databases, graph databases, and multi-agent frameworks.

Tools in this posting

  • Azure
  • Google Cloud (GCP)
  • Kubernetes
  • Terraform
Source — Tool mentions in context
- Solid understanding of AI infrastructure patterns, including model training, inference, model serving, RAG architectures, data pipelines, and emerging agentic AI frameworks. - Experience deploying and supporting hybrid cloud infrastructure, with working knowledge of Azure, GCP, or similar cloud platforms. - Strong technical foundation in Linux, networking, storage, distributed systems, observability, performance engineering, and production operations.
- Drive the design, deployment, and optimization of GPU-based infrastructure supporting model training, inference, retrieval-augmented generation (RAG), agentic AI, data pipelines, and other production AI workloads across on-premises, hybrid, and cloud environments. - Serve as the senior technical leader for AI infrastructure, guiding decisions related to Kubernetes, distributed systems, GPU orchestration, infrastructure automation, observability, performance engineering, capacity planning, and operational resilience. - Develop and mature enterprise platform capabilities including AI model serving, scalable data and compute infrastructure, vector and graph database ecosystems, AI orchestration frameworks, and microservices architectures.
- Deep expertise designing and operating AI, machine learning, GPU, HPC, or distributed computing platforms in production environments. - Strong experience with Kubernetes, containerized platforms, GPU orchestration technologies, and modern infrastructure automation practices. - Expertise with NVIDIA GPU ecosystems, including GPU lifecycle management, workload optimization, and large-scale compute environments.
- 10+ years of experience in a related field with a 6-8 years’ experience managing staff - Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline; advanced STEM degrees or relevant cloud, Kubernetes, AI, or infrastructure certifications are a plus.
- Strong technical foundation in Linux, networking, storage, distributed systems, observability, performance engineering, and production operations. - Experience leveraging Infrastructure as Code, CI/CD, and automation tools such as Terraform, Ansible, Helm, ArgoCD, GitLab, Jenkins, or similar technologies. - Ability to influence senior stakeholders and effectively communicate complex technical concepts to both engineering and executive audiences.

Job description

View original posting ↗

AI & High Performance Computing (HPC) Infrastructure Engineering Manager, Senior Director

We’re seeking a future team member for the role of AI & High Performance Computing Infrastructure Engineering Manager, Senior Director to join our Enterprise Infrastructure Delivery organization. This role is in New York, NY.

The AI and HPC Infrastructure Engineering Manager will lead the engineering and strategic evolution of the bank’s AI, machine learning, and high-performance computing infrastructure platforms. Responsible for managing a team of AI Infrastructure Engineers resources who design, operate, automate, and scale GPU-based infrastructure supporting model training, inference, agentic AI, data pipelines, and production AI workloads.

 In this role, you'll make an impact in the following ways

  • Lead and develop a team of AI Infrastructure Engineers responsible for building, operating, and scaling the firm's AI, machine learning, and high-performance computing (HPC) platforms.
  • Define and execute the strategic roadmap for AI infrastructure, partnering across Engineering, Architecture, Security, Risk, Compliance, Production Services, and AI/ML teams to deliver secure, scalable, and resilient platforms.
  • Drive the design, deployment, and optimization of GPU-based infrastructure supporting model training, inference, retrieval-augmented generation (RAG), agentic AI, data pipelines, and other production AI workloads across on-premises, hybrid, and cloud environments.
  • Serve as the senior technical leader for AI infrastructure, guiding decisions related to Kubernetes, distributed systems, GPU orchestration, infrastructure automation, observability, performance engineering, capacity planning, and operational resilience.
  • Develop and mature enterprise platform capabilities including AI model serving, scalable data and compute infrastructure, vector and graph database ecosystems, AI orchestration frameworks, and microservices architectures.
  • Partner closely with AI/ML engineering teams to accelerate model deployment, improve platform adoption, optimize performance, and enable the delivery of innovative AI solutions.
  • Advance automation across infrastructure provisioning, configuration management, monitoring, incident response, workload onboarding, and platform lifecycle management.
  • Establish and govern operating models, service standards, production readiness requirements, support processes, and reliability objectives to ensure exceptional platform availability and user experience.
  • Manage strategic technology vendor relationships, influencing product roadmaps, driving technical evaluations and proof-of-concepts, and ensuring alignment with business objectives.
  • Communicate platform strategy, investment priorities, operational performance, risks, and roadmap progress to executive leadership and technology governance forums.

 

To be successful in this role, we're seeking the following

  • Proven leadership experience managing infrastructure engineering, platform engineering, DevOps, SRE, or production operations teams in large-scale enterprise environments.
  • Deep expertise designing and operating AI, machine learning, GPU, HPC, or distributed computing platforms in production environments.
  • Strong experience with Kubernetes, containerized platforms, GPU orchestration technologies, and modern infrastructure automation practices.
  • Expertise with NVIDIA GPU ecosystems, including GPU lifecycle management, workload optimization, and large-scale compute environments.
  • Solid understanding of AI infrastructure patterns, including model training, inference, model serving, RAG architectures, data pipelines, and emerging agentic AI frameworks.
  • Experience deploying and supporting hybrid cloud infrastructure, with working knowledge of Azure, GCP, or similar cloud platforms.
  • Strong technical foundation in Linux, networking, storage, distributed systems, observability, performance engineering, and production operations.
  • Experience leveraging Infrastructure as Code, CI/CD, and automation tools such as Terraform, Ansible, Helm, ArgoCD, GitLab, Jenkins, or similar technologies.
  • Ability to influence senior stakeholders and effectively communicate complex technical concepts to both engineering and executive audiences.
  • Demonstrated success developing technology strategies, operating models, governance frameworks, and long-term infrastructure roadmaps.
  • Strong vendor management experience, including strategic partnerships, architecture reviews, technical assessments, and escalation management.

 

Preferred Qualifications

  • Experience within financial services or other highly regulated industries.
  • Expertise supporting AI platforms subject to stringent security, compliance, audit, and operational resilience requirements.
  • Experience with large-scale GPU environments, DGX platforms, SuperPOD architectures, InfiniBand, high-performance storage, or enterprise AI infrastructure.
  • Familiarity with modern AI serving and orchestration technologies such as Triton, vLLM, NVIDIA NIM, KServe, Seldon, LangGraph, vector databases, graph databases, and multi-agent frameworks.
  • Experience with capacity planning, usage analytics, cost optimization, GPU quota management, and chargeback/showback models.
  • 10+ years of experience in a related field with a 6-8 years’ experience managing staff 
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline; advanced STEM degrees or relevant cloud, Kubernetes, AI, or infrastructure certifications are a plus.

Your next step

  • Have your CV and examples of relevant work ready.
  • Check the listed location, eligibility and core experience before starting.
  • Ask the employer about the salary range before committing time to the process.

Complete your application on eofe.fa.us2.oraclecloud.com. The employer’s form will show what is required.

Already applied? Track this application

Source & posting history

View original posting ↗

Source notes

Source excerpts

Selected passages from the saved posting. Check the full description for conditions and exceptions.

Pay

No pay amount identified in the saved description.

Location & working pattern

New York, NY, United States

- Define and execute the strategic roadmap for AI infrastructure, partnering across Engineering, Architecture, Security, Risk, Compliance, Production Services, and AI/ML teams to deliver secure, scalable, and resilient platforms. - Drive the design, deployment, and optimization of GPU-based infrastructure supporting model training, inference, retrieval-augmented generation (RAG), agentic AI, data pipelines, and other production AI workloads across on-premises, hybrid, and cloud environments. - Serve as the senior technical leader for AI infrastructure, guiding decisions related to Kubernetes, distributed systems, GPU orchestration, infrastructure automation, observability, performance engineering, capacity planning, and operational resilience.
More source context
- Solid understanding of AI infrastructure patterns, including model training, inference, model serving, RAG architectures, data pipelines, and emerging agentic AI frameworks. - Experience deploying and supporting hybrid cloud infrastructure, with working knowledge of Azure, GCP, or similar cloud platforms. - Strong technical foundation in Linux, networking, storage, distributed systems, observability, performance engineering, and production operations.
Work authorization

No clear work-authorization passage found. Eligibility is unconfirmed.

Status in our records
Active
First seen by us
Sep 12, 2026
Recorded sightings
17
Last seen by us
Oct 7, 2026
Employer says posted
Sep 10, 2026

These dates show when we found the listing. Check the employer’s website to confirm it is still accepting applications.

Report an error

See how this role fits your experience

Add your resume to compare the role’s scope, tools and requirements with your experience.

Find answers in the posting

AI
How answers work

AI selects complete passages from this posting. Check them for conditions and exceptions.

Uses this posting and your question. No profile needed.