ParallelIQ
Careers

Build the optimization layer for the next generation of AI.

We're a small team solving a hard infrastructure problem. If you've run GPU workloads at scale and know the pain firsthand, we'd love to talk.

We ship to real clusters.

Every line of code runs on production GPU infrastructure. No toy benchmarks, no sandboxed demos — we operate where it matters.

Operators stay in the loop.

We build tools that augment human judgment, not replace it. Every recommendation is reviewable, reversible, and auditable.

Clarity over cleverness.

We explain what our system sees and why it recommends what it recommends. No black boxes — inside or outside the product.

Open roles

Founding Engineer — AI Infrastructure Optimization

EngineeringRemote (US)Full-time · Competitive compensation + equity
Apply

Most GPU clusters running AI inference are wasting 30–50% of their compute — not because of bad hardware, but because nobody has built a system that understands what's actually running inside the cluster well enough to do something about it. That's what we're building at Paralleliq. This is a founding-team role. You'll work directly with the founder on the scanner, the optimization rules engine, and the remediation workflow layer. The core challenge is not building another monitoring dashboard — it's building a system that can distinguish between a GPU at 12% utilization that means dark capacity, healthy speculative decoding, a network bottleneck, or thermal throttling, and act correctly on each.

Responsibilities

  • Extend the piqc scanner to collect model-aware telemetry from vLLM, SGLang, DCGM, Kubernetes, and network operators
  • Develop optimization and fault-detection rules that map inference failure patterns — misplacement, KV cache pressure, network saturation, thermal throttling — to concrete remediation steps
  • Design fact schemas that correctly represent inference topology: disaggregated prefill/decode, LMCache, MIG/MPS partitioning, speculative decoding
  • Build and maintain the control plane, rules evaluation pipeline, and Temporal-based remediation workflows
  • Work directly with customers on on-prem deployments, constraint configuration, and rule validation against live cluster telemetry
  • Own the full stack from telemetry collection through the operator dashboard

Qualifications

  • Deep familiarity with vLLM internals — continuous batching, PagedAttention, KV cache management, sequence length behavior, the V1 engine
  • Working knowledge of inference ecosystem technologies: LMCache, disaggregated prefill/decode (llmd, Dynamo), speculative decoding, MIG/MPS/time-slicing, SGLang
  • Observability experience — Prometheus, DCGM GPU telemetry, Kubernetes events, and how inference frameworks export metrics and where those metrics mislead
  • The ability to reason from a cluster state to a root cause and a remediation — given a set of signals, you can identify what's wrong and why
  • 5+ years of backend or platform engineering experience in Python (FastAPI, SQLAlchemy) with strong Kubernetes proficiency
  • Experience running or building tooling for GPU workloads in production environments
  • Comfort working across the full stack; backend and inference domain are the priority
  • Startup operating mode — you move fast, write for the reader, and push back when something doesn't make sense

Solutions Engineer

Customer EngineeringRemote (US)Full-time
Apply

Work directly with GPU cloud providers and inference platform teams to onboard, deploy, and get value from Paralleliq. You're the bridge between product and customer.

Responsibilities

  • Lead technical onboarding for new customers — from cluster access to first recommendation surfaced
  • Diagnose GPU waste patterns in customer environments and translate findings into actionable insights
  • Work with the engineering team to close product gaps discovered during customer deployments
  • Build repeatable onboarding playbooks and technical documentation
  • Run discovery calls and technical demos with prospects at GPU cloud providers and enterprise AI teams

Qualifications

  • 3+ years in a solutions engineering, customer engineering, or technical account management role
  • Hands-on experience with Kubernetes and cloud infrastructure (AWS, GCP, or Azure)
  • Familiarity with AI/ML infrastructure — model serving, GPU utilization, inference optimization
  • Ability to read and understand Python and YAML; light scripting for customer environments
  • Strong communicator — equally comfortable in a Slack thread and a C-suite demo
  • Experience working with early-stage products where the playbook doesn't yet exist

Apply

For the engineering role, tell us what you've built or operated in AI inference infrastructure, and describe one specific failure mode you've debugged in a GPU cluster — what you saw, what it turned out to be, and how you figured it out. Don't see your role? Select “General Interest” and tell us what you've built.

Get more from the cluster you already have.

Start for Free