Teleport brand banner

System Reliability Engineer / T2 Support Engineer

TeleportGurugram, IndiaPosted Apr 22
via Workable

About the Role

We are looking for an engineer who enjoys understanding how systems behave in real production, not just writing features. This role is responsible for maintaining reliability, stability, and smooth functioning of our live platform running on Google Cloud.

You will act as the first technical owner of production systems — monitoring services, investigating alerts, resolving issues, and performing controlled configuration and operational changes. This role works closely with backend developers, QA, and infrastructure teams to prevent incidents and reduce downtime.

This is not a call-center support role and not a pure development role — it is a hands-on technical position focused on debugging, incident handling, and system operations.

Tech Stack

  • Google Cloud Platform (Compute, Logging, Monitoring)
  • Java (Spring Boot based microservices)
  • MongoDB
  • Apache Kafka (event-driven architecture)
  • Redis cache
  • Linux servers

Key Responsibilities

Production Monitoring & Alert Handling

  • Monitor application health, latency, errors, consumer lag, database connections, and resource utilization
  • Acknowledge and investigate monitoring alerts
  • Perform first-level troubleshooting and stabilize services
  • Identify whether issue is infra, application, database, or messaging related

Incident Response

  • Participate in on-call rotation
  • Diagnose production incidents and restore services with minimal downtime
  • Safely restart services, scale instances, or rollback deployments when required
  • Communicate incident status to stakeholders

Technical Support & Operational Changes

  • Handle technical support tickets requiring engineering understanding
  • Update configurations and feature flags
  • Manage scheduled jobs / cron triggers
  • Trigger or replay events in Kafka
  • Assist in minor Java configuration/code fixes when needed
  • Coordinate production releases

Database & Messaging Operations

  • Investigate MongoDB performance issues and slow queries
  • Monitor and resolve Kafka consumer lag and stuck messages
  • Manage Redis cache behavior (TTL, eviction, connection issues)

Logs & RCA

  • Analyze logs and metrics to determine root cause of issues
  • Prepare basic Root Cause Analysis (RCA) reports
  • Suggest preventive actions to reduce recurring incidents

Required Skills

Core Technical Skills

  • Good understanding of Linux commands and server behavior
  • Experience analyzing application logs and debugging runtime issues
  • Basic Java knowledge (stack trace reading, configuration changes, rebuild & deploy)
  • Practical experience with MongoDB (indexes, connections, slow queries)
  • Understanding of Kafka concepts (consumer, offset, lag, partitions)
  • Basic Redis knowledge (caching behavior, TTL)

Cloud & Tools

  • Hands-on experience with any cloud platform (GCP preferred / AWS acceptable)
  • Experience using monitoring tools (GCP Monitoring, Prometheus, Grafana, ELK, or similar)
  • Understanding of REST APIs and HTTP status codes

What We Expect From You

  • Ability to investigate problems logically rather than randomly restarting services
  • Comfort working with live production systems
  • Willingness to participate in on-call support
  • Strong ownership mindset and attention to detail
  • Good communication during incidents

Good to Have

  • Experience in e-commerce, fintech, logistics, or high-traffic systems
  • Exposure to CI/CD pipelines and deployments
  • Basic scripting (Shell or Python)
  • Experience writing RCA documents

Experience

3 – 6 years of relevant experience in production support, application support, SRE, DevOps operations, or similar roles.

Why Join Us

  • Direct exposure to real distributed systems
  • Hands-on production debugging experience
  • Opportunity to learn system architecture deeply
  • Close interaction with development and platform teams

Important Note

This role involves handling live production systems and occasional on-call responsibilities. Candidates interested only in feature development or pure infrastructure automation may not find this role suitable.

Similar roles

  • Tacto logo

    Value Engineer

    Tacto·Munich, Germany

    YOUR IMPACT Procurement is the single biggest cost position for industrial companies, often 60-70% of revenue. Tacto brings AI to strategic procurement so teams can move faster, become more resilient and unlock savings at scale. However, winning a customer takes more than a good platform. It takes a business case that convinces the C-level (primarily CEO and CFO). As our…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • Navflex logo

    Senior Manager, Robotic Platforms Engineering - M/F/D

    Navflex·Munich, Germany

    About Navflex At Navflex, we're pioneering the future of logistics automation through cutting-edge AI and robotics. Our autonomous mobile robots (AMRs) are transforming the way goods are loaded and unloaded, enabling plug-and-play solutions that streamline operations and enhance efficiency across the global supply chain for some of the world’s most demanding warehouse Environments. Our international, cross‑disciplinary team in the EU…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • Navflex logo

    Senior Functional Safety Engineer/Safety Systems Engineer - M/F/D

    Navflex·Munich, Germany

    Non‑negotiables for this role * 3+ years hands-on programming of SICK safety controllers using SICK Safety Designer (professional, production use). * Electrical Engineering background with strong hands-on commissioning / troubleshooting on real vehicles. * Confident Python programming and ability to read and debug existing Python code (integration of safety system with application logic). ABOUT NAVFLEX Navflex builds autonomous mobile robots…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • Michael Wessel logo

    IT System Engineer Microsoft Cloud & Infrastructure (m/w/d)

    Michael Wessel·Hannover, Germany

    IT System Engineer Microsoft Cloud & Infrastructure (m/w/d) Du möchtest Microsoft-Infrastrukturen nicht nur betreiben, sondern planen, modernisieren und weiterentwickeln? Dann verstärke unser Team Modern Workspace & Infrastructure. Wir suchen sowohl erfahrene System Engineers und Consultants als auch technisch starke Nachwuchskräfte, die sich gezielt in die Microsoft-Cloud- und Infrastrukturwelt entwickeln möchten. Entscheidend ist für uns nicht, dass du bereits jedes Produkt…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • MaxAccelerate logo

    Senior AI & Solutions Engineer

    MaxAccelerate·Dubai, United Arab Emirates

    SENIOR AI & SOLUTIONS ENGINEER AI • SALESFORCE • LLMS • INTELLIGENT CRM • FULL-STACK ENGINEERING BUILD THE FUTURE OF AI-POWERED ENTERPRISE SOLUTIONS We are looking for an exceptional Senior AI & Solutions Engineer to join our team and help us push the boundaries of what is possible with AI, Salesforce and next-generation CRM technology. This is not a traditional…

    • Remote
    • Full-time
    • Green Visaself-sponsored, no tie
  • Wehrtyou logo

    Research Engineer

    Wehrtyou·United States

    <p>Hudson River Trading (HRT) is a quantitative trading firm at the forefront of technological innovation. We build and deploy cutting-edge systems within one of the world’s most advanced computing environments to power our global trading operations. Our non-siloed, collaborative coding environment empowers talented engineers to make significant contributions and see their impact daily. At HRT, you'll be challenged to solve…

    • On-site
    • Full-time
  • SandboxAQ logo

    ML Research Engineer, AI for Life Sciences

    SandboxAQ·United Kingdom

    ABOUT SANDBOXAQ SandboxAQ is a high-growth company delivering AI solutions that address some of the world's greatest challenges. The company’s Large Quantitative Models (LQMs) power advances in life sciences, financial services, navigation, cybersecurity, and other sectors. We are a global team that is tech-focused and includes experts in AI, chemistry, cybersecurity, physics, mathematics, medicine, engineering, and other specialties. The company…

    • On-site
    • Full-time
  • 360Learning logo

    Solutions Engineer DACH (m/f/d)

    360Learning·Berlin, Germany

    As a Solutions Engineer, you will be working at the intersection of Sales, Product and Customer Success, working with top learning and development organizations to align their growth objectives to 360learning capabilities. As the first Solutions Engineer for the German market, you will play a crucial role in ensuring the region’s success, but you will have the support from the…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo