AI Engineer - Product Engineer

TraversalApplyPublished 2 months agoFirst seen 2 hours ago
Apply

About Traversal

Traversal is the AI Site Reliability Engineer (AI SRE) for the enterprise.

Production complexity was already outpacing what engineering teams could manage manually, and AI-generated code is accelerating that gap. Traversal is built for that challenge, autonomously understanding and reasoning across even the largest, most complex production environments to diagnose, fix, and prevent incidents. Our mission is to free engineers from endless firefighting and give them more time to focus on creative, high-impact work.

Today, Traversal operates in mission-critical environments at some of the world’s largest enterprises. Our roots remain deeply embedded in AI research, and we’ve brought together researchers from institutions including MIT, Harvard, Berkeley, Columbia, and Cornell with world-class technical staff and operators from companies like Google, Meta, Datadog, ServiceNow, and Citadel Securities to take on one of the hardest problems for AI to solve. Traversal is backed by Sequoia Capital, Kleiner Perkins, Hanabi, NFDG, and American Express Ventures.

The Role

Our core product is our AI SRE agent. As a Product Engineer on the Agents team, you'll own how that agent behaves, where it shows up, and whether it holds up under real production load. That means working end-to-end on the harness and context engineering that shape agent reasoning, the streaming and orchestration layers that put it in front of engineers mid-incident, and the observability and evaluation systems that tell us whether any of it is actually working.

Three areas drive impact on this team:

  • Agent Behavior: Improving how the agent reasons, decides, and communicates through harness design and context engineering. This is prompt and tool architecture, not prompt tweaking: what the agent sees, when it acts, what it does with ambiguity, and how we measure whether a change made it better.
  • Live Surface Area: Getting agent intelligence into the places engineers already work, in real time. Slack channels, the web app, and whatever comes next. Streaming, orchestration, and long-running work that survives restarts and reconnects.
  • Stability at Scale: Our agents run for hours against noisy, high-volume customer telemetry. Making that reliable means observability into agent behavior, scalable data patterns, and systems that degrade gracefully instead of silently.

Each of these requires product vision and autonomy. You'll be deciding what to build as often as you're building it.

Responsibilities

  • Agent Development: Build AI agents with a mind for user experience and capabilities like long-horizon work, proactiveness, resilience, readability, and evidence citation.
  • User Experience: Innovate and drive the evolution of AI UX/UI design, rapidly iterating based on user feedback to improve overall usability.
  • Harness and Context Engineering: Design the tool interfaces, context assembly, and decision logic that determine how well the agent performs on real incidents.
  • Evaluation: Build the eval harnesses and offline datasets that let us ship agent changes with confidence rather than vibes.
  • Live Delivery: Design and implement the streaming, orchestration, and API layers that carry agent output into Slack and the product in real time.
  • Data and Infrastructure: Own efficient storage, retrieval, and processing across PostgreSQL, Redis, Kafka, and S3, in support of agents operating on large-scale time-series and topological data.
  • Product Impact: Translate agent capability into features that reduce cognitive load for on-call engineers, iterating quickly on user feedback.

Requirements

  • 3+ years of software engineering experience, with a strong focus on full-stack or backend systems.
  • Hands-on experience building with agentic AI systems or LLM-powered products, including familiarity with tool use, context management, and the failure modes that come with both.
  • Strong experience with Python and web frameworks such as FastAPI.
  • Comfort working in a React and TypeScript codebase, enough to ship a feature end to end without waiting on someone else for the frontend.
  • Experience deploying applications on AWS, working with ECS or Kubernetes, Postgres for data storage, and S3 for large-scale object storage.
  • Proven ability to lead in fast-paced startup environments, with limited resources, shifting priorities, and minimal structure.

You don't need to meet every requirement. If you have deep experience in a subset of these areas, we encourage you to apply.

Nice to Have

  • Experience with durable execution or workflow orchestration frameworks such as Temporal.
  • Background in observability, monitoring, or incident response tooling, as a builder or a heavy user.
  • Experience running evals for LLM systems, or otherwise bringing measurement discipline to non-deterministic software.
  • Background in large-scale, complex, data-driven applications.

Compensation

We offer competitive compensation, startup equity, health insurance, and additional benefits. The U.S. base salary range for this full-time, in-person role in New York is $150,000–$300,000, plus equity and benefits. Our salary ranges are based on location, level, and role. Individual compensation is determined by experience, skills, and job-related knowledge.

Why You Should Join Us

Traversal is a place to take on hard, meaningful problems with real ownership from day one. You’ll work alongside people who challenge you to grow, learn constantly, and help define a new category of infrastructure software. We think long term, move quickly, and hold a high bar without taking ourselves too seriously.

We offer competitive salary and equity packages, health insurance, fertility benefits, a great tech setup stipend and flexible time off. Plus in-office snacks, team happy hours and outings, an annual company offsite, and plenty of built in time to collaborate across teams.

Traversal is fully in-office, 5 days a week, based in New York near Madison Square Park. We have a collaborative, hard-working culture and are energized by building the future of AI-powered software maintenance.