Job Title :Support Engineer and Incident Management
Location: Costa Mesa, CA (Onsite)
Schedule: 24/7 Operations / Shift-Based Support
Duration: 12+ Months
Role Summary:
- We are seeking an Incident Management (IM) Support Engineer to support 24/7 operations for connected vehicle and telematics systems.
- This role focuses on monitoring system health, identifying anomalies, managing production incidents, coordinating incident response, and ensuring timely resolution within established SLAs.
- The engineer will be responsible for end-to-end incident handling, including anomaly detection, incident triage, bridge initiation, stakeholder communication, incident coordination, resolution tracking, and post-incident activities.
Key Responsibilities:
-
Monitor the health and availability of production systems supporting connected vehicle and telematics environments.
-
Monitor application, infrastructure, API, database, and service health using tools such as:
-
Datadog
-
Grafana
-
Dynatrace
-
MaxGauge
-
Elastic/log monitoring tools
-
-
Identify abnormal system behavior, performance degradation, errors, latency, service failures, and potential outages.
-
Analyze dashboards, alerts, metrics, logs, and application behavior to determine incident scope and impact.
-
Perform proactive monitoring to identify issues before they impact customers.
-
Validate alerts and distinguish genuine production incidents from false positives or monitoring noise.
-
Handle incidents across different priority levels, including P1, P2, P3, and P4.
-
Initiate and manage incident bridges/war rooms for high-priority incidents.
-
Coordinate application, infrastructure, network, cloud, vendor, and other technical teams during incidents.
-
Act as an Incident Commander during major incidents and drive the incident toward resolution.
-
Maintain accurate incident timelines, action items, decisions, and communication records.
-
Ensure timely stakeholder communications throughout the incident lifecycle.
-
Track and maintain incident management SLAs, including MTTD, MTTA, and MTTR.
-
Facilitate post-incident reviews and Root Cause Analysis (RCA).
-
Document contributing causes and corrective/preventive actions.
-
Maintain incident management dashboards, KPIs, and operational reports.
-
Identify recurring issues and work with engineering/product teams on reliability improvements.
-
Maintain and improve incident management procedures, playbooks, runbooks, and operational processes.
-
Participate in incident simulations, readiness exercises, and operational reviews.
Domain Expertise – Telematics & Connected Vehicles:
Candidates should have strong experience in the Telematics / Connected Vehicle domain, including:
-
Telematics Control Unit (TCU)
-
eSIM and OTA provisioning
-
Backend telematics platforms
-
Data flows between vehicles, cloud platforms, and mobile applications
-
IoT systems and messaging technologies
-
Messaging queues
-
Bulk provisioning
-
Notification systems
-
API gateways
-
Load balancers
-
Mobile and web applications
-
Remote vehicle services such as:
-
Remote lock/unlock
-
Remote start/stop
-
Charging control
-
Climate pre-conditioning
-
-
Geo-services and connected vehicle features
-
Identity and authentication services
-
Device management
-
CAN bus signals
-
Firmware and OTA update processes
-
Regional/compliance considerations
Experience coordinating incidents across multiple technical teams, vendors, cloud providers, and support organizations is highly valuable.
Required Qualifications:
-
5+ years of experience in Operations or Service Management.
-
4+ years of Incident Management experience in large-scale, 24/7 production environments.
-
Demonstrated experience managing P1/P2 incidents and leading incident bridges.
-
Strong Incident Commander experience with the ability to make decisions under pressure.
-
Strong background in Telematics / Connected Vehicle systems.
-
Experience with vehicle remote services and connected vehicle technologies.
-
Hands-on experience with monitoring and observability tools such as:
-
Datadog
-
Dynatrace
-
Grafana
-
MaxGauge
-
-
Strong understanding of the complete Incident Management lifecycle:
Detection → Triage → Severity Assignment → Bridge Initiation → Stakeholder Communication → Resolution → RCA -
Working knowledge of ITIL Incident, Problem, Change, and Service Level Management processes.
-
Experience with tools such as:
-
Jira
-
Confluence
-
Jenkins
-
xMatters
-
ServiceNow or similar ticketing platforms
-
-
Strong written and verbal communication skills.
-
Strong stakeholder management and coordination skills.
-
Ability to analyze telemetry, logs, metrics, and time-series data to develop root-cause hypotheses.
-
Experience working with SLAs/SLOs and operational KPIs such as MTTD, MTTA, MTTR, incident volume, recurrence, and customer impact.
Tools & Technologies:
Monitoring & Observability
-
Datadog
-
Dynatrace
-
Grafana
-
MaxGauge
-
Elastic / log monitoring
Incident & Collaboration
-
Jira
-
ServiceNow
-
Confluence
-
Teams / Zoom
-
xMatters
Data & Monitoring
-
Time-series analysis
-
Log aggregation
-
Distributed tracing
-
Metrics
-
Alerting
-
Noise reduction
Education & Experience:
-
Bachelor's degree in Engineering, Computer Science, Information Technology, or a related field is preferred.
-
Equivalent practical experience may be considered.