Are you from India? 🇮🇳
👉 Check Today's Deals on Amazon IndiaAbout You
We are seeking a technically curious and detail-oriented Operations Engineer to join our Global Technical Operations (GTO) team. The ideal candidate will thrive in a fast-paced, collaborative environment and be excited to monitor and investigate production issues globally. You’ll help enhance incident detection and response, analyze trends in production data, and improve communication with partners and stakeholders during incidents.
Key Qualifications
Strong troubleshooting skills, experience with observability platforms, and scripting knowledge are crucial. A background in SRE, DevOps, production operations, or NOC environments supporting high-availability platforms (payments, e-commerce, SaaS, or gaming) is preferred. You must communicate clearly in English—both in writing and verbally—when updating incidents, providing shift handoffs, and engaging in status page communications.
If you’re passionate about maintaining critical systems and continuously improving operational processes while driving incident resolution for game developers and players worldwide, we would love to hear from you!
Role: Operations Engineer in Kuala Lumpur
**Location:** Kuala Lumpur
**Company:** Xsolla
About Us
Xsolla is a global commerce company offering robust tools and services to help developers navigate the challenges of the video game industry. From indie to AAA titles, companies partner with Xsolla to fund, distribute, market, and monetize their games. We believe in the future of gaming and are committed to creating opportunities for creators. Headquartered in Los Angeles, Xsolla operates as the merchant of record and has supported over 1,500 game developers in reaching more players and growing their businesses globally. With innovative paths to profits, developers have everything they need to succeed.
For more information, visit [xsolla.com](https://xsolla.com).
Responsibilities
– **Dashboard Monitoring:** Serve as the primary monitor for the GTO Operational Dashboard in Datadog, detecting anomalies by correlating signals across APM, logs, metrics, and synthetic tests. Determine if alerts require an incident ticket or can be resolved through immediate investigation.
– **Incident Investigation:** Create incident tickets in JIRA Service Management, conduct initial technical investigations using Datadog, determine blast radius and likely root causes, and route incidents to the appropriate teams.
– **End-to-End Incident Management:** Own lower-severity incidents from detection to resolution, executing runbook procedures and escalating issues when necessary.
– **Major Incident Support:** Assist the TSO Lead during major incidents by providing real-time data and maintaining incident tickets with live entries.
– **Communication:** Draft incident communications, including internal updates and customer-facing status page reports under the TSO Lead’s direction.
– **Incident Analysis:** Analyze incident trends and production bugs during non-incident periods, compiling data to identify patterns and contribute findings to product and engineering reports.
– **Health Reports:** Publish periodic health reports for critical applications and track Post-Incident Review action items.
– **Automation and Documentation:** Build operational automation tools and contribute to runbook development to ensure procedures are repeatable.
– **Shift Handoffs:** Conduct structured shift handoffs covering active incidents and upcoming deployments while participating in knowledge transfer sessions with SREs.
– **Leadership Coverage:** Step in for the TSO Lead during absences, managing escalation decisions and stakeholder communications.
Qualifications
– 4+ years of experience in SRE, DevOps, production operations, NOC, or technical operations in environments with high availability, preferably in payments, e-commerce, SaaS, or gaming.
– **Troubleshooting Skills:** Strong investigative skills to trace alerts and user-reported symptoms through application logs, infrastructure metrics, and network paths.
– **Observability Platform Experience:** Hands-on experience with Datadog or similar platforms is crucial for monitoring and alert management.
– **Scripting Proficiency:** Familiarity with at least one scripting language, such as Python, Go, or Bash, is essential for automation and tool building.
– **Communication Skills:** Excellent written and verbal communication skills in English are necessary for clear incident documentation and updates.
– **Cloud Infrastructure Knowledge:** Understanding of Kubernetes and cloud platforms (GCP preferred) is required for effective incident investigation.
– **Understanding of SLOs:** Knowledge of SLOs and incident management tools is important, along with experience in AI/ML-assisted operations.
Nice to Have
– Experience in the gaming, payments, or fintech industries, particularly in environments with strict uptime requirements.
– Familiarity with database operations and CI/CD pipelines.
– ITIL Foundation certification is a plus.
Compensation and Benefits
– **Salary Range:** RM144,000 – RM216,000 per year.
– **Work Tools:** Latest Mac workstations and essential hardware to maximize effectiveness.
– **Professional Growth:** Opportunities for free training and participation in specialized conferences.
– **Health Coverage:** Comprehensive health insurance for employees and dependents.
– **Flexible Work Hours:** Organize your day based on personal needs and team demands.
– **Work Environment:** Comfortable office space with no dress code.
The responsibilities of this position may evolve over time for organizational success. By submitting your application, you consent to Xsolla conducting background checks as permitted by law post-interview. Xsolla KL Sdn Bhd prioritizes your privacy and does not sell or distribute your data. For inquiries about data privacy, please contact [email protected].
For more vacancies, visit [Careers](https://xsolla.com/careers).
Location
Kuala Lumpur
Source link
