Senior Forward Deployed Solution Engineer (Poland)
Who We Are
Spectro Cloud helps platform teams and cloud providers modernize and manage infrastructure for the AI era without adding more tools or operational complexity.
With PaletteAI, enterprises, public sector organizations, neoclouds and sovereign clouds can build, govern and operate full-stack environments across VMs, Kubernetes, edge, regulated and air-gapped locations, and AI infrastructure. PaletteAI Launchpads help teams start quickly with urgent outcomes such as VMware migration, token cost control or edge modernization, then scale into enterprise-wide lifecycle management, governance and fleet operations on the same platform.
We are looking for a Sr. Forward Deployed Solution Engineer (FDSE) who is equally comfortable writing automation, troubleshooting a production Kubernetes cluster, whiteboarding a GPU fabric, and working with Customer leaders to be the tip of our FDE practice and guide the Customer to successful technical and business outcomes.
What You'll Do
You will join the FDE team and at the point with our customer where business or mission goals meet technical reality. You will translate desired outcomes into deployable architectures, build the automation and integrations needed to make those architectures real, and guide each deployment from first working capability through production readiness, scaled adoption, and operational handoff and working closely with our other FDE practice resources and the aligned Deployment Strategist.
This is a hands-on engineering not traditional staff augmentation and not a role where the work ends with a slide deck or demo. You will work across Palette / PaletteAI, Kubernetes, cloud, bare metal, edge, GPU systems, data pipelines, high-performance networking, and the NVIDIA software and hardware ecosystem. You will also bring what you learn back into Product and Engineering so that the next deployment is faster, more standardized, and easier for our FDE teams and Customers to operate independently.
About You
You thrive where the problem is important, the environment is complex, and the path is not yet fully defined. You can move between code, infrastructure, networking, user workflows, architectural reviews and guidance with other FDE team members, and Customer conversations without losing the thread of the outcome.
You are curious without being careless, confident without pretending to know everything, and pragmatic without accepting permanent shortcuts. You communicate early when risk changes, write down what others will need later, and know that the best technical solution is one the customer can operate successfully. You enjoy ambiguity, learn quickly, communicate directly, and care deeply about the difference between technology that is installed and capability that is genuinely useful.
We do not expect one person to have deployed every product and technology outlined in this description. We do expect deep Kubernetes and automation capability, strong systems reasoning, customer-facing judgment, and the ability to become productive quickly in unfamiliar parts of the stack.
Minimum Qualifications
- Bachelor’s degree in Computer Science, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience.
- 6+ years of professional experience in software engineering, platform engineering, site reliability engineering, DevOps, infrastructure engineering, solutions engineering, or a comparable customer-facing technical role.
- Production experience designing, deploying, upgrading, and troubleshooting Kubernetes clusters and cloud-native applications.
- Strong Linux and distributed-systems fundamentals, including processes, networking, storage, identity, certificates, resource management, and failure analysis.
- Proficiency in at least one modern programming language preferably Go or Python and the ability to produce maintainable, tested, production-quality code.
- Hands-on experience with infrastructure as code and configuration automation using technologies such as Terraform, Ansible, Cluster API, Helm, Kustomize, and equivalent tools.
- Experience implementing CI/CD or GitOps workflows using platforms such as GitHub Actions, GitLab CI, Jenkins, Argo CD, Flux, or comparable systems.
- Working knowledge of Kubernetes security and multi-tenancy, including RBAC, namespaces, network policy, admission controls, quotas, secrets, pod security, and tenant isolation.
- Experience with at least one major public cloud and at least one datacenter deployment model such as bare metal, VMware, KubeVirt, Proxmox, or OpenStack.
- Working knowledge of GPU-enabled infrastructure, including the relationship among firmware, drivers, CUDA, container runtimes, Kubernetes scheduling, device plugins, telemetry, and workload frameworks.
- Working knowledge of the AI data and model lifecycle from data ingestion through model deployment, inference, monitoring, and feedback.
- Solid networking fundamentals, including TCP/IP, routing, switching, DNS, BGP, load balancing, overlays, MTU, latency, bandwidth, congestion, and packet-level troubleshooting.
- Familiarity with high-performance AI networking concepts such as RDMA, RoCE, InfiniBand, SR-IOV, GPUDirect, and collective communication.
- Ability to communicate architecture, tradeoffs, status, and risk clearly to engineers, operators, executives, and nontechnical stakeholders.
- Demonstrated ability to work through ambiguous requirements, prioritize effectively, and deliver iterative value in a fast-moving environment.
- Ability to work onsite with Customer for discovery, deployment, testing, and operational milestones.
- Authorization to work in the stated location and ability to satisfy customer-specific access requirements.
Preferred Qualifications
- 8+ years of relevant experience, including ownership of complex production deployments or strategic customer programs.
- Direct experience with Spectro Cloud Palette / PaletteAI, Cluster API, Kubernetes fleet management, edge Kubernetes, or declarative full-stack lifecycle platforms.
- Current CKA, CKS, CKAD, relevant cloud, networking, security, or NVIDIA certification.
- Production experience designing or operating NVIDIA Cloud Partner (NCP) or comparable multi-tenant GPU cloud environments.
- Deep experience with tenant isolation across bare metal, virtual machines, Kubernetes control planes, GPU resources, storage, and high-performance networks.
- Hands-on experience with NVIDIA GPU Operator, Network Operator, DCGM, MIG, vGPU, NIM, Triton, TensorRT-LLM, vLLM, NCCL, or NVIDIA AI Enterprise.
- Experience with GPU schedulers and AI platform technologies such as Run:ai, Slurm, Volcano, KAI, Ray, Kubeflow, ClearML, or comparable systems.
- Experience with distributed training or high-throughput inference on H100, H200, B200, GB200, GB300, or other modern accelerator platforms.
- Hands-on design or operational experience with NVIDIA Spectrum-X Ethernet fabrics, Quantum InfiniBand, ConnectX adapters, BlueField DPUs or SuperNICs, and DOCA OR other high-speed, low-latency networking technologies.
- Experience configuring or troubleshooting GPUDirect RDMA, GPUDirect Storage, RDMA, RoCE, PFC, ECN, ECMP, SR-IOV, Multus, InfiniBand, InfiniBand partitions, and topology-aware scheduling.
- Experience building or operating SONiC or Cumulus Linux leaf-spine fabrics, including FRR, BGP, EVPN/VXLAN, VRF, zero-touch provisioning, automated validation, upgrades, and telemetry.
- Experience integrating BlueField in DPU or restricted modes for network, security, storage, isolation, or telemetry offload; familiarity with DPU provisioning, BMC, Redfish, attestation, and lifecycle management.
- Experience with high-performance storage and AI data platforms, including object storage, NVMe-oF, parallel file systems, data lakes, streaming, feature stores, vector databases, model registries, or metadata and lineage systems.
- Experience in disconnected, air-gapped, sovereign, regulated, safety-sensitive, or classified environments.
- A record of contributing reusable software, open-source projects, reference architectures, technical publications, or customer enablement materials.
Location: This role will be in Wroclaw, Poland and requires a daily onsite presence of 5 days a week to meet project security compliance.
This position requires access to NATO classified information. The successful candidate must possess, or be eligible to successfully undergo, the vetting process for a NATO Secret security clearance.