Job Title :Support Engineer and Incident Management 
Location: Costa Mesa, CA (Onsite)
Schedule: 24/7 Operations / Shift-Based Support
Duration: 12+ Months

Role Summary:

  • We are seeking an Incident Management (IM) Support Engineer to support 24/7 operations for connected vehicle and telematics systems.
  • This role focuses on monitoring system health, identifying anomalies, managing production incidents, coordinating incident response, and ensuring timely resolution within established SLAs.
  • The engineer will be responsible for end-to-end incident handling, including anomaly detection, incident triage, bridge initiation, stakeholder communication, incident coordination, resolution tracking, and post-incident activities.

Key Responsibilities:

  • Monitor the health and availability of production systems supporting connected vehicle and telematics environments.

  • Monitor application, infrastructure, API, database, and service health using tools such as:

    • Datadog

    • Grafana

    • Dynatrace

    • MaxGauge

    • Elastic/log monitoring tools

  • Identify abnormal system behavior, performance degradation, errors, latency, service failures, and potential outages.

  • Analyze dashboards, alerts, metrics, logs, and application behavior to determine incident scope and impact.

  • Perform proactive monitoring to identify issues before they impact customers.

  • Validate alerts and distinguish genuine production incidents from false positives or monitoring noise.

  • Handle incidents across different priority levels, including P1, P2, P3, and P4.

  • Initiate and manage incident bridges/war rooms for high-priority incidents.

  • Coordinate application, infrastructure, network, cloud, vendor, and other technical teams during incidents.

  • Act as an Incident Commander during major incidents and drive the incident toward resolution.

  • Maintain accurate incident timelines, action items, decisions, and communication records.

  • Ensure timely stakeholder communications throughout the incident lifecycle.

  • Track and maintain incident management SLAs, including MTTD, MTTA, and MTTR.

  • Facilitate post-incident reviews and Root Cause Analysis (RCA).

  • Document contributing causes and corrective/preventive actions.

  • Maintain incident management dashboards, KPIs, and operational reports.

  • Identify recurring issues and work with engineering/product teams on reliability improvements.

  • Maintain and improve incident management procedures, playbooks, runbooks, and operational processes.

  • Participate in incident simulations, readiness exercises, and operational reviews.

Domain Expertise – Telematics & Connected Vehicles:

Candidates should have strong experience in the Telematics / Connected Vehicle domain, including:

  • Telematics Control Unit (TCU)

  • eSIM and OTA provisioning

  • Backend telematics platforms

  • Data flows between vehicles, cloud platforms, and mobile applications

  • IoT systems and messaging technologies

  • Messaging queues

  • Bulk provisioning

  • Notification systems

  • API gateways

  • Load balancers

  • Mobile and web applications

  • Remote vehicle services such as:

    • Remote lock/unlock

    • Remote start/stop

    • Charging control

    • Climate pre-conditioning

  • Geo-services and connected vehicle features

  • Identity and authentication services

  • Device management

  • CAN bus signals

  • Firmware and OTA update processes

  • Regional/compliance considerations

Experience coordinating incidents across multiple technical teams, vendors, cloud providers, and support organizations is highly valuable.

Required Qualifications:

  • 5+ years of experience in Operations or Service Management.

  • 4+ years of Incident Management experience in large-scale, 24/7 production environments.

  • Demonstrated experience managing P1/P2 incidents and leading incident bridges.

  • Strong Incident Commander experience with the ability to make decisions under pressure.

  • Strong background in Telematics / Connected Vehicle systems.

  • Experience with vehicle remote services and connected vehicle technologies.

  • Hands-on experience with monitoring and observability tools such as:

    • Datadog

    • Dynatrace

    • Grafana

    • MaxGauge

  • Strong understanding of the complete Incident Management lifecycle:
    Detection → Triage → Severity Assignment → Bridge Initiation → Stakeholder Communication → Resolution → RCA

  • Working knowledge of ITIL Incident, Problem, Change, and Service Level Management processes.

  • Experience with tools such as:

    • Jira

    • Confluence

    • Jenkins

    • xMatters

    • ServiceNow or similar ticketing platforms

  • Strong written and verbal communication skills.

  • Strong stakeholder management and coordination skills.

  • Ability to analyze telemetry, logs, metrics, and time-series data to develop root-cause hypotheses.

  • Experience working with SLAs/SLOs and operational KPIs such as MTTD, MTTA, MTTR, incident volume, recurrence, and customer impact.

Tools & Technologies:

Monitoring & Observability

  • Datadog

  • Dynatrace

  • Grafana

  • MaxGauge

  • Elastic / log monitoring

Incident & Collaboration

  • Jira

  • ServiceNow

  • Confluence

  • Teams / Zoom

  • xMatters

Data & Monitoring

  • Time-series analysis

  • Log aggregation

  • Distributed tracing

  • Metrics

  • Alerting

  • Noise reduction

Education & Experience:

  • Bachelor's degree in Engineering, Computer Science, Information Technology, or a related field is preferred.

  • Equivalent practical experience may be considered.

Apply
Come join the Ninja family!

Fill out the form and one of our expert recruiters will get back to you.