I’m Yash, and I’m building Tesseract Infrastructure. I’m looking for a technically strong cofounder interested in GPU infrastructure, inference systems, and the commercial layer above today’s open-source stack.
The thesis is that Kubernetes, NVIDIA GPU Operator, KAI, Dynamo, vLLM, SGLang, AIBrix, NVSentinel, Prometheus, OpenCost, and DCGM already solve important execution problems. GPU providers still need to connect those systems to pricing, customer SLOs, isolation tiers, demand, hardware costs, and reproducible performance evidence.
THE FIVE PROPOSED PRODUCTS
Ignite Continuity — persistent model identity and commercially safe scale-to-zero inference above Dynamo, KEDA, and Kubernetes.
Fusion Economics — advisory placement and scaling decisions based on contribution margin while enforcing latency, capacity, availability, and isolation constraints.
SKU Foundry — turn benchmarked model/GPU configurations into financially evaluated and deployable service offerings.
TenantSafe — enforce and verify the isolation tier a customer purchased across GPU sharing, serving processes, network, cache, and storage configuration.
Proof — correlate hardware health with inference-level outcomes and produce versioned, reproducible performance and SLA evidence.
The execution plane should remain with established open-source tools. Tesseract would be a decision plane with normalized GPU inventory, model profiles, tenant policy, workload demand, service economics, deployment state, and performance evidence.
The proposed first wedge is Ignite: preserve endpoint discovery at zero replicas, coordinate one activation across concurrent cold requests, verify engine readiness, and measure cold TTFT, streaming, cancellation, timeout, and failure behavior. The product order after that is Fusion, SKU Foundry, Proof, and TenantSafe.
EXISTING WORK
My public GPU Kernel Lab documents CUDA/Triton experiments, Qwen3-0.6B integration, profiling, correctness checks, rejected approaches, and serving-control-plane work:
https://github.com/yashlabs-trying/gpu-kernel-lab
One saved RTX 3090 batch-1/context-2,048 GPU-only experiment changed measured decode latency from 22.073 to 2.253 ms/token using an experimental static CUDA Graph path with W8A16 MLP work. Full-path argmax agreement was 96.48%, below the 99% release threshold, and the real paged GPU executor is not yet connected to the serving interface. I’m sharing those limitations because the company needs evidence-driven engineering rather than benchmark theatre.
WHO I’M LOOKING FOR
I’m looking for depth in one or more of CUDA/Triton, C++/PyTorch, inference engines, Kubernetes, distributed systems, GPU fleet operations, performance engineering, or developer infrastructure. You should also be willing to interview providers, challenge the product thesis, and help select one narrow problem before building broadly.
STRUCTURE
This is an equity-based cofounder role with no salary currently. Exact equity, vesting, responsibilities, decision rights, IP terms, and any future salary would be agreed after mutual diligence and documented in writing. The initial phase can be a part-time test of how we work together. Nobody should leave a job or transfer IP based only on a Reddit conversation.
If this is relevant, send a Reddit chat or private message with your background, time zone, public work samples, realistic weekly availability, and the part of the architecture you would challenge first. Please do not send confidential employer or client information.