Collaboration.Ai logo

Senior AI Engineer - Agentic Systems & Data Pipelines

Collaboration.AiWorldwidePosted 6h agoremote
via Findwork

Who We Are

Collaboration.Ai is a mission-focused, AI-powered software and services company based in Minnesota, with employees, partners, and customers around the world. We unite people, technology, and purpose to accelerate breakthroughs that transform industries, empower communities, and create a more sustainable future. We collaborate with organizations across the defense ecosystem, helping them navigate complex challenges and drive transformative change.

Our Products

NetworkOS — NetworkOS is an AI-powered platform that aligns people, purpose, ideas, and expertise in real-time, generating actionable insights to propel movements forward.

CrowdVector — CrowdVector is an integrated solution marketplace and innovation management platform that rapidly uncovers new ideas and advances breakthroughs to fuel movements.

To learn more about us, visit collaboration.ai.

About the Role

You'll build the agentic systems and data pipelines behind NetworkOS's AI capabilities: production agent workflows built on industry-leading agent SDKs and harnesses, MCP servers, and Agent Skills standards; the eval and observability layer that keeps LLM quality measurable; and the ingestion pipelines that turn messy, diverse data sources into queryable knowledge.

This is an execution seat, not an ivory tower. You'll commit code every week, ship agents as product capability rather than demos, and help shape a roadmap that's heading deep into graph + agents territory — for customers in defense, public sector, and regulated enterprise.

Agents in production. Pipelines that hold. Evals that keep everyone honest.

What You'll Do

  • Ship production agent systems — design, build, and operate agentic workflows (agent SDKs, MCP servers, Agent Skills standards) powering AI-driven matching, analysis, and data intelligence

  • Operationalize LLM quality — build the eval and observability layer with Langfuse, golden datasets, LLM-as-judge patterns, and FinOps-style tracking so every workflow has measurable quality, cost, and latency

  • Engineer data pipelines — robust ingestion of documents, structured data, and external sources into searchable knowledge bases with quality validation, deduplication, and incremental updates

  • Own retrieval quality — hybrid search combining vector, keyword, and metadata retrieval, continuously improved through reranking, query expansion, and contextual compression

  • Accelerate with AI — build custom MCP tools and Agent Skills that make the whole engineering team measurably faster

  • Execute alongside the team — pair with full-stack engineers on AI integration points, contribute to incident response for AI services, and keep your hands in the code

Our Tech Stack

  • Languages: Python (primary); Kotlin (core platform language at CAI); TypeScript/Node.js and other modern languages (secondary)

  • AI/ML: FastAPI, Pydantic; multi-provider LLM SDKs (Anthropic, OpenAI, and others)

  • Agentic Tooling: Claude Code/Codex/etc.; industry-leading agent SDKs and harnesses; MCP servers; Agent Skills standards

  • LLM Operations: Langfuse + evals (golden datasets, LLM-as-judge); in-house FinOps tracking (token usage, latency, cost); multi-provider orchestration including AWS Bedrock

  • Search & Retrieval: Vector databases, OpenSearch, embedding models

  • Data: PostgreSQL, Amazon S3; streaming pipelines (Kafka/Kinesis) where needed

  • Infrastructure: Docker, Kubernetes (AWS EKS); DataDog + OpenTelemetry observability

What We're Looking For

Must-Haves

  • 7+ years of professional software engineering experience, with 3+ years focused on AI/ML or data engineering

  • Production agentic/LLM application experience — built and operated systems around LLM APIs (Anthropic, OpenAI) serving real users: agents, tool-use, or orchestrated LLM workflows

  • Data engineering background — robust, scalable pipelines for AI/ML workloads

  • LLM operations experience — evals and observability for production LLM systems (quality, cost, latency)

  • Production retrieval experience — vector databases and/or search engines (OpenSearch, Elasticsearch)

  • Modern Python stack proficiency — FastAPI, Pydantic, async/await, modern dependency management

  • AI-native workflows — demonstrated ability to leverage Claude Code/Codex or similar agentic coding tools to accelerate development

  • Experience with Docker, Kubernetes, and AWS

  • US citizenship required (DoD contracting — IL4/IL5 environments — and FedRAMP compliance)

Nice-to-Haves

  • Deep agentic ecosystem experience — Agent Skills standards, custom MCP servers, agent SDKs across major vendors

  • Advanced RAG expertise — GraphRAG, agentic RAG, contextual retrieval, reranking strategies

  • Graph data experience — knowledge graphs, graph databases, or graph-based retrieval

  • Model selection & rightsizing — matching models to domain-specific use cases across quality, cost, and latency tradeoffs

  • Streaming data experience (Kafka, Kinesis) for real-time knowledge base updates

  • Research background, open-source contributions, or an advanced degree in ML/IR/NLP

Why Join Collaboration AI?

Real AI engineering, not a wrapper shop. Production agents, hybrid retrieval, continuous evals, and a roadmap heading into graph + agents — with the autonomy to shape how it's built.

AI-native by default. We build with AI, not just for AI. Agentic coding tools (Claude Code/Codex/etc.), agent SDKs and harnesses, MCP servers, and Agent Skills standards are how we work daily — you'll both use and build them.

Work that matters. Defense, public sector, and regulated industries — SOC 2 and NIST compliance, FedRAMP readiness, and customers whose missions demand AI they can trust.

Small, senior team. Early-stage impact with your work visible from week one. You'll help set the bar for how AI engineering is done here.

 

Skills

  • llm
  • python
  • claude
  • rag
  • kotlin
  • anthropic
  • node
  • kafka
  • ml
  • agents
  • nlp
  • opentelemetry
  • s3
  • datadog
  • fastapi
  • postgresql
  • eks
  • elasticsearch
  • aws
  • docker

Similar roles

  • Edfinity logo

    Senior Software Engineer, remote

    Edfinity·Worldwide

    About Edfinity Edfinity is the category leader in courseware and assessment technology for higher-ed STEM. We're NSF-supported, bootstrapped, and built by a close-knit remote team of senior engineers. We're growing fast and leading the field in AI-powered courseware that moves the needle on educator productivity and student outcomes. Your work will reach hundreds of institutions, and you will help decide…

    • Remote
    • Full-time
  • AutoGPT logo

    GTM Lead

    AutoGPT·Worldwide

    Remote · ~3 days/week · 12-week contract · £500–£650/day · Immediate start · Paid working trial for finalists ABOUT AUTOGPT At AutoGPT, we're building toward a future where AI agents do real work for people — not just answer questions, but take action, complete tasks, and move work forward. Our mission is to democratize AI by making powerful digital assistants…

    • Remote
    • Contract
  • Stripe logo

    Program Manager, Performance and Talent Planning

    Stripe·Worldwide

    WHO WE ARE ABOUT STRIPE Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow revenue, and accelerate new business opportunities. Our mission is to increase the GDP of the internet, and we have a significant amount of work ahead. That means you have an…

    • Full-time
  • Cloudflare logo

    Senior Customer Engineer, Named

    Cloudflare·Worldwide

    About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s largest networks that powers millions of websites and other Internet properties for customers ranging from individual bloggers to SMBs to Fortune 500 companies. Cloudflare protects and accelerates any Internet application online without adding hardware, installing software, or…

    • Full-time
  • Spotlight PA / Grist logo

    Environment reporter

    Spotlight PA / Grist·Worldwide

    Summary: Spotlight PA and Grist are seeking an environment reporter to cover the impacts of climate change in Pennsylvania. The ideal candidate will be interested in exploring the intersectionality of a warming planet with, well, everything — from land use, business, energy, water, and policy decisions to environmental justice and health — as well as solutions to the climate crisis.…

    • Remote
    • Full-time
  • Toptal logo

    AI/ML Engineer for an AI-Driven E-Commerce Platform

    Toptal·Worldwide

    We are seeking an AI Engineer to develop AI-first applications and integrate custom models and infrastructure into an existing e-commerce platform. The role focuses on building agent-led product experiences, including internal automation, dynamic site customization, and customer-facing conversational commerce workflows. General information This engagement supports an e-commerce startup building a platform around conversational design and intelligent automation. The current objective…

    • Remote
    • Full-time
  • Spotify logo

    Senior Product Quality Analyst - AI Voice

    Spotify·London, Canada

    The Personalization team makes deciding what to play next easier and more enjoyable for every listener. From Blend to Discover Weekly, we’re behind some of Spotify’s most-loved features. We built them by understanding the world of music and podcasts better than anyone else. Join us and you’ll keep millions of users listening by making great recommendations to each and every…

    • Remote
    • Full-time
    • Express EntryPR day one, no employer
  • Samsara logo

    Staff Software Engineer

    Samsara·Worldwide

    About the role: Samsara (NYSE: IOT) sits at the center of hardware, software, AI, and the physical world. The platform processes 25+ trillion data points annually from IoT sensors, cameras, and connected devices across thousands of organizations, and that dataset is compounding: it grew from 4 trillion to 25+ trillion in four years. We're still early in turning that data…

    • Remote
    • Full-time