
We’re seeking Founding AI Engineer to join Full-time onsite in San Francisco, CA 3 – 5 years of experience in applied AI/ML engineering, ideally computer vision or multimodal systems. Location : San Francisco, CA On-site work policy: On-site 5 days per week in San Francisco. Full-time position Salary: $180K – $240K Hiring Count: Looking to hire 2 candidate for this role Tech stack: Python, PyTorch, vLLM, Triton, Ray Serve, ONNX, TensorRT, Hugging Face Transformers, LangChain, RAG, RLHF, SFT, Quantization (GPTQ, AWQ), Edge AI, Multimodal LLMs, Vision-Language Models, Docker Visa sponsorship available: For H-1B transfers and TN visas supported. No new H-1B sponsorship. Candidates must currently reside in the USA or Canada. About This Role We’re looking for an AI Engineer with 1-5 years of experience in applied AI/ML who has shipped agentic multimodal systems to real users – not just RAG wrappers or single-turn chatbots. You should be comfortable building production VLM pipelines with real hardware constraints (latency, power, connectivity) and have a track record of owning AI systems end-to-end in fast-paced environments. Bonus points if you have wearable AI, autonomous driving, or industrial domain experience. What you’ll do:
- Build and ship the production agentic-VLM pipeline running on industrial smart glasses – multi-step, tool-using visual-reasoning loops against real customer workflows (SOPs, inspection, field service)
- Own model orchestration and runtime optimization for edge inference, trading off model quality against latency with graceful fallback across connectivity conditions
- Design and build the eval harness and data flywheel from scratch – failure-mode capture, customer-data fine-tune loops, and measurable model quality improvements each release
- Ship real-time voice-video AI interfaces adapted to different end-user profiles: video-heavy, conversational speech, and proactive alerts
- Build RAG pipelines for efficient creation and querying of enterprise knowledge bases from field operator data
- Drive multimodal model training when needed for on-premise deployments: open-source model SFT, RL post-training, and quantization
Role requirements
- Seniority – 3 – 5 years of experience in applied AI/ML engineering, ideally computer vision or multimodal systems.
- Work experience – Shipped multimodal and computer vision systems in the VLM era. (production, not demos, not pure research).
- Owned the model layer end to end – Experience at a startup OR AI team building production AI products, preferably as a founder/CTO/founding engineer.
- Big- company tenure (Meta Reality Labs, Snap, Apple Vision, Google, big streaming infra) is a BONUS only when on a directly- relevant team (AR/smart- glasses, real- time video/streaming, on- device/edge ML) OR paired with a builder signal (founder / early- startup / side projects / OSS).
- A long single big- tech tenure with no builder signal and no relevant- team work is a NEGATIVE.
- Production AR / wearable AI experience (Meta Reality Labs, Snap, Apple Vision) or autonomous driving Computer Vision.
- Industrial domain exposure (data centers, energy grid, aerospace, manufacturing).
- Education – Strong CS/ML/Eng background OR demonstrated equivalent shipping record.
- Master’s with vision or multimodal research component.
- Hard skills Applied VLM / multimodal engineering: makes VLMs reliably do visual reasoning in production (image/video understanding, detection/segmentation as needed) in the modern VLM/VLA era. Depth is in shipping, hardening, and applied fine- tuning, not pretraining from scratch – this means VISION- language / video- language multimodality (images, video, VLM/VLA), NOT sensor- fusion, materials, audio- only, or time- series ‘multimodal’. Applied agentic AI / model orchestration experience (vs. pure research) Evals discipline: rigorous eval harnesses (ground- truth, trajectory and tool- call accuracy, regression) to compare models/orchestrations and drive iteration On- prem / self- hosted model deployment: serving and optimizing open- weight ML models on customer hardware (for high- IP environments). Hands- on fine- tuning and deployment of a vLLM is a proxy In- context grounding / RAG against a knowledge base, with tool- and- KB wiring (a real plus; FDEs own the customer- specific integration) Soft skills Balances core AI depth with applied product mindset. Genuinely mission- aligned with CV/wearables/industrial AI.
- Miscellaneous – Ideally willing to join a hacker house (live on site). At least willing to work on- site 5 days/week in SF. – We are not looking Pure researcher focused on pretraining / training from scratch with no shipped product – Classical CV vocabulary only (Faster R- CNN, IOU segmentation) without VLM and agentic awareness. – We are not looking Pure Big- tech- only profile without ownership of a shipped AI product. – ‘Multimodal’ that means sensor/materials/signal fusion or audio- only/time- series rather than vision- language. $5,000 Referral Bonus for successful placements!
#J-18808-Ljbffr