Customer Success Engineer (HQ2)

BlitzyApplyPublished 2 months agoFirst seen 3 hours ago
Apply

Blitzy is a Cambridge, MA based AI software development platform on a mission to revolutionize the software development life cycle by autonomously building custom software to unlock the next industrial revolution. We're transforming how enterprises build software, turning enterprise requirements into enterprise grade code with an agentic software development platform that can autonomously execute 80% of the quantum of software development work. We're backed by multiple tier 1 investors, and have proven success as founders of previous start-ups.

Our Culture

Who we are:

Led by two pioneering co-founders we are one of the fastest growing companies in the U.S., creating our own category of enterprise autonomous software development. We automate thousands of hours of software development for our customers, which includes strong representation within the Fortune 500.

How we work:

  • We move Blitzy Fast: Time is both our company’s and our clients’ most precious asset. We move quickly and decisively to innovate internally and deliver exceptional software externally.
  • Championship Mindset: We operate like a professional sports team. We win as a team by holding ourselves and each other to high standards, collaborating in-person, and remaining focused on the mission.
  • Passion for Invention: We’re pushing the frontier of what’s possible, requiring constant innovation and iteration.
  • We Work for the Customer: We focus on delivering outsized value to the customers we work with and expanding those relationships into deep, meaningful partnerships.
  • We believe in being ‘everyday athletes’: taking care of ourselves so we can bring our best minds to work. We promote great sleep, movement, and restorative activities for optimal mental performance. It makes for a happier and more productive team.

About the Role

This role plays a critical role in supporting our clients and maintaining a stable, reliable environment throughout the full lifecycle- from installation and upgrades to day-to-day operations. You’ll work closely with L1 Support to triage and resolve technical issues, while partnering with Engineering to escalate and troubleshoot unresolved defects. The role spans Kubernetes, Docker, and major cloud platforms, with a strong focus on customer support, technical problem-solving, and operational reliability.

Responsibilities

  • Deploy and install the platform into customer environments, and troubleshoot installation issues.
  • Support ongoing upgrades and day-to-day operation, keeping customer environments stable.
  • Work alongside L1 to triage and resolve customer-reported issues, driving them to resolution or escalation.
  • Diagnose failures across the stack: compute, networking, storage, and the services running on it.
  • Reproduce issues safely against live (often multi-tenant) environments using read-only diagnostics first.
  • Build and maintain dashboards, monitors, and runbooks so recurring issues get faster to fix: or stop recurring.
  • Write up clear, evidence-backed escalations and post-incident notes.
  • Communicate status and resolution to customers clearly and on time.

Qualifications

  • Experience debugging distributed systems across services, queues, databases, and networks using evidence-based troubleshooting.
  • Experience with Kubernetes, Docker, and major cloud platforms (GCP, AWS, or Azure), including managed Kubernetes, logging, IAM, and storage.
  • Strong monitoring and observability skills with tools such as Datadog, including logs, traces, dashboards, alerts, and request correlation.
  • Experience troubleshooting networking and WebSockets, including DNS, routing, TLS, NAT, and load-balancer issues.
  • Working knowledge of Python, Redis, message queues, SQL, and PostgreSQL.
  • Experience with GitHub, GitHub Enterprise, Azure DevOps, and/or GitLab.
  • Familiarity with CI/CD, Helm, and container deployment pipelines; ArgoCD is a plus
  • Ability to securely manage secrets, credentials, and certificates; Vault is strongly preferred.
  • Comfortable troubleshooting both Linux and Windows environments.
  • Experience with incident management and ticketing tools such as Jira is a plus.
  • Customer-facing support, SRE, or on-call background is a plus.
  • Methodical, evidence-first approach to troubleshooting with a focus on validating root causes and minimizing risk.
  • Strong understanding of multi-tenant environments, access controls, and potential blast radius.

Hours & On-Call:

This is a customer support role, and the hours can be unconventional. Customers operate primarily in US time zones, so coverage is anchored to US business hours (roughly ET–PT). If you're based outside the US, expect your working day to shift accordingly.

Incidents don't keep office hours. Expect a rotating on-call schedule and occasional evening, early-morning, or weekend escalations outside a standard 9–5. We structure for it: rotations are shared fairly, on-call is compensated/time-off-in-lieu per policy, and we protect recovery time after heavy incidents.

If you're not comfortable with US-aligned hours and periodic off-hours on-call, this likely isn't the right role, and that's completely fine.

U.S. GenAI startup, Pune Office, Private Limited Company

Full Time Employment directly to our Private Limited Company. We are committed to an enduring and robust presence in Pune. Our goal is to enable you to have a long enduring career at Blitzy with opportunity for advancement throughout your career at the company. If you want a role you can turn into a career, read on! If you are looking for a short term arrangement, we recommend you look elsewhere!

Blitzy is an equal opportunity employer committed to building a diverse and inclusive team. We believe different perspectives make us stronger.