Remote job
Remote confirmedSenior Solutions Architect – AI Infrastructure Security (remote in Europe)
Mirantis
Key points from the posting
- Tech stack
- k0rdent AIKubernetesKubeVirtLinuxTPMRedfishSR-IOVKMSHSMMIGvGPUTransformer
- Seniority:
- Senior
Automatically extracted
Our assessment
- Our reading of the full posting text confirms it: fully remote.
- The posting states no salary. Comparable roles in our index (10 postings): median 3,850 euros per month, middle range 2,815 to 6,050 euros.
- 14 more open roles from this employer in our index. 13 of them fully remote.
Automated assessment by nomado24, not an employer statement.
Job description
About the role
We are seeking a Senior Solutions Architect with deep security expertise to join the Voyager team in the Mirantis Office of the CTO. You will own the security solutioning of k0rdent AI, our platform for building and operating GPU clouds and AI factories, from metal to model.
You will define how k0rdent AI environments are secured end to end: the hardware and firmware they boot from, the hosts, virtual machines and Kubernetes clusters they run, the networks and storage they share, the secrets and certificates that hold them together, and the AI models and applications they serve. You will turn that design into published reference architectures, working code and proofs of concept that run on real hardware with our customers and partners.
As part of the Office of the CTO, research is a core part of the job. You will explore threats, designs and technologies before customers ask for them, prototype ideas that may not ship, and challenge established practice when a better approach exists. Your work will shape where k0rdent AI goes next, not only how it is deployed today.
This role combines research with hands-on, outward-facing work. You will split your time between exploring new designs, writing architecture, building and validating it in the lab, and explaining it to engineers, security teams, executives and conference audiences. You will work with Solutions Architects, Partner Management and Engineering colleagues across many countries and time zones.
Key Responsibilities
Reference architecture
- Design and publish security reference architectures and solution designs for k0rdent AI, covering single-tenant private AI clouds, multi-tenant GPU clouds and hybrid deployments that span public cloud and on-premises infrastructure.
- Define a defence-in-depth model across the stack, from hardware root of trust and firmware to hosts, virtual machines, Kubernetes, networks, storage, models and applications.
- Define tenant isolation boundaries and their guarantees, and document where each boundary is strong, where it is weak, and what it costs to strengthen.
- Map designs to the controls and frameworks customers must meet, and document trade-offs (risk, cost, performance, operability) with defensible reasoning a customer can follow.
Platform and infrastructure security
- Bare-metal and host security: secure and measured boot, TPM-based attestation, BMC and Redfish hardening, firmware integrity, and hardened Linux hosts.
- Virtual machine security: hypervisor and KubeVirt hardening, VM isolation, and the security implications of PCI passthrough and SR-IOV for GPUs and NICs.
- Kubernetes security: RBAC, Pod Security Standards, admission control and policy-as-code, runtime security, and secure multi-cluster and multi-tenant operation.
- Secrets management: secret stores, KMS and HSM integration, envelope encryption, secret delivery to workloads, and rotation.
- Certificates and PKI: internal certificate authorities, automated issuance and rotation, mTLS between services, and workload identity.
- Network security: segmentation and zero-trust designs, Kubernetes network policy, tenant isolation on front-end and GPU fabrics, and DPU-based security enforcement.
- Storage security: encryption at rest and in transit, key management, and tenant isolation on shared block, file and object storage used for training data and model weights.
- Public cloud security: secure landing zones, identity and access management, and network controls on major public clouds, and how they extend to hybrid deployments.
AI model and application security
- Define how Transformer-based models are protected through their lifecycle: provenance and signing of model weights, integrity of training and fine-tuning data, and secure model registries.
- Secure inference services and AI applications against threats such as prompt injection, data and model exfiltration, model theft and abuse of agentic tool access.
- Design isolation for shared GPU infrastructure (MIG, vGPU, passthrough) and assess confidential computing options for protecting models and data in use.
- Secure the software supply chain for platforms, containers and models: image signing, SBOMs, vulnerability management and provenance.
Research and exploration
- Track and evaluate emerging security technologies, standards and threats, such as confidential computing on GPUs, remote attestation, post-quantum cryptography, model signing and AI agent security.
- Prototype alternative designs in the lab, including ones customers have not yet asked for, and measure them against current practice.
- Challenge established or customer-preferred designs when evidence points to a better option, and make the case with data.
- Publish findings as internal research notes and design proposals and, where appropriate, as external papers, blog posts or talks.
Code and proofs of concept
- Build and run proofs of concept with customers and partners, on Mirantis lab …
This role is provided by an external source. Applications are handled on the source website.
