Engineer the backbone of high-growth SaaS & AI platforms.

We are a boutique, senior-only platform engineering practice. No bureaucracy, no junior handoffs, and no legacy maintenance traps. Work alongside seasoned infrastructure architects solving high-impact problems across AWS, Kubernetes, and GPU inference.

Why Engineers Join Illusio

🌍

100% Remote-First

Work from wherever you are most productive. We operate asynchronously with clear documentation and minimal meetings.

🧠

Senior Peer Group

Collaborate directly with engineers averaging 8+ years of production experience, including CNCF Kubestronauts and AWS Solution Architects.

Zero Busywork

No manual ticket queues or legacy patch firefighting. We design greenfield platforms, automate FinOps, and modernize architectures.

🎓

Continuous Mastery

Unlimited budget for AWS, CKA/CKS certifications, personal sandbox AWS environments, and high-impact technical conferences.

Open Positions

3 Active Roles

Senior Platform Engineer (AWS & Kubernetes)

Full-time 100% Remote 6+ Years Exp Kubernetes / EKS
Apply Now

We are looking for a seasoned Platform Engineer to design, build, and optimize internal developer platforms (IDPs) and production Amazon EKS clusters for fast-scaling SaaS companies. You will lead client platform architecture, automate GitOps workflows, and implement intelligent autoscaling with Karpenter.

What You Will Do

  • Design and deploy resilient, multi-tenant Kubernetes clusters on AWS EKS using Terraform and Helm.
  • Implement declarative GitOps delivery pipelines using ArgoCD and automated drift detection.
  • Migrate legacy architectures and manual node groups to Karpenter dynamic bin-packing and Spot orchestration.
  • Enforce production security baselines: EKS Pod Identity (IRSA), Cilium/Calico network policies, and OPA/Kyverno admission controllers.
  • Collaborate directly with client VP of Engineering and Staff Engineers as a trusted technical advisor.

Requirements

  • 6+ years of hands-on platform or infrastructure engineering experience with AWS and Kubernetes in production.
  • Deep mastery of Terraform (reusable modules, state locking, multi-account structures) and Helm.
  • Strong networking fundamentals (VPC peering, Transit Gateway, Route 53, CoreDNS, Cilium/eBPF or AWS VPC CNI).
  • Track record of owning zero-downtime cluster upgrades and mission-critical production migrations.
  • CKA, CKAD, or CKS certifications are strongly preferred.
Reference: REQ-PLT-2026 Apply via email →

Senior DevOps & Site Reliability Engineer (AWS)

Full-time 100% Remote 6+ Years Exp AWS & Terraform
Apply Now

Join as a Senior DevOps & SRE to build bulletproof infrastructure pipelines, automate multi-account governance via AWS Control Tower, and maintain four-nines reliability across client production fleets.

What You Will Do

  • Build automated CI/CD golden paths using GitHub Actions, GitLab CI, and container scanning tools (Trivy, Grype).
  • Implement comprehensive observability stacks: Prometheus, Grafana, Loki/OpenSearch, and AWS CloudWatch telemetry.
  • Establish SLO/SLI error budget frameworks and automated incident triage playbooks.
  • Conduct AWS Well-Architected Reviews and execute hands-on CloudSpend cost remediation (gp3 conversions, rightsizing, NAT gateway optimization).
  • Automate security and compliance guardrails aligning with SOC 2 Type II and CIS benchmarks.

Requirements

  • 6+ years in DevOps, SRE, or Cloud Systems Engineering in high-availability cloud environments.
  • Expertise in AWS core services (EC2, VPC, IAM, RDS, S3, CloudFront, Route53, KMS).
  • Proficiency in Infrastructure as Code (Terraform) and at least one scripting language (Python, Bash, or Go).
  • Strong Linux systems internals troubleshooting skills (cgroups, systemd, memory/CPU profiling, network socket debugging).
  • AWS Solutions Architect Professional or DevOps Professional certification preferred.
Reference: REQ-OPS-2026 Apply via email →

AI Inference & Platform Infrastructure Engineer

Full-time 100% Remote 5+ Years Exp GPU & vLLM / Triton
Apply Now

Architect and scale the next generation of generative AI inference clusters. You will optimize LLM serving latency, implement GPU autoscaling on Kubernetes, and pioneer custom AWS Inferentia/Trainium deployments for AI clients.

What You Will Do

  • Deploy and tune high-throughput inference runtimes (vLLM, NVIDIA Triton, TensorRT-LLM, TGI) on EKS.
  • Engineer fast model weight loading mechanisms to reduce cold start times from minutes to seconds.
  • Implement Karpenter GPU autoscaling, node consolidation, and MIG (Multi-Instance GPU) slicing.
  • Instrument AI FinOps telemetry: tracking cost-per-million-tokens, KV cache hit rates, and GPU memory saturation.
  • Evaluate and benchmark transformer inference on AWS Inferentia2 (Inf2) vs. NVIDIA Ada/Hopper architectures.

Requirements

  • 5+ years of software/systems engineering with at least 2+ years focused on ML/AI inference or GPU systems.
  • Hands-on experience deploying open-weights models (Llama, Mistral, Whisper, embedding models) in production.
  • Strong understanding of CUDA runtimes, GPU device plugins in Kubernetes, and NVIDIA DCGM metrics.
  • Experience with Kubernetes container networking and high-performance storage (EBS gp3, io2, JuiceFS, NVMe).
  • Proficiency in Python and systems languages (Go or C++).
Reference: REQ-AIML-2026 Apply via email →

Don't see an exact match?

We are always eager to meet exceptional senior cloud architects, FinOps specialists, and Kubernetes engineers. Send us your GitHub, LinkedIn, or CV, and let's explore if there's a mutual fit.

Email careers@theillusio.com Get in Touch