[北京/上海] [30-60k] 特斯拉热招资深SRE工程师
关于我们
官方网站:www.tesla.cn/
公司规模:2 万 人
公司地址:中国(上海)自由贸易试验区临港新片区江山路 5000 号
关于我们
The Role
As a Senior Site Reliability Engineer at Tesla China, you will own reliability, observability, automation, and operational efficiency for critical business systems and infrastructure. You will keep services stable, secure, and scalable under production load and peak traffic; manage monitoring, change, incidents, and performance in an engineering-driven way; drive SLO/SLI adoption; partner with development, platform, and business teams on end-to-end reliability; and continuously reduce incident impact and manual toil through automation and data.
Responsibilities
- Own stability operations for in-scope systems and platforms; build and continuously improve monitoring, alerting, logging, and tracing to speed detection and localization
- Define and track SLO/SLI/SLA and error budgets; use data to surface risk and drive improvements in availability, latency, capacity, and change quality
- Contribute to high-availability architecture and disaster recovery; land backup/restore, degradation and rate limiting, fault isolation, and emergency drills
- Manage configuration, deployment, scaling, inspection, and day-to-day operations through code and automation to reduce repetitive manual work
- Establish and improve incident response; cover on-call and major incidents; complete root cause analysis (RCA) and closed-loop remediation
- Partner with development, platform, network, security, and business teams to identify launch and architecture risks early and improve releasability and rollback safety
- (Splunk focus) Own Splunk platform architecture, deployment, data onboarding, parsing, and knowledge object management; build logging and observability capabilities for troubleshooting, audit, security, and business analytics
- Produce and maintain runbooks, emergency playbooks, and technical documentation; codify standards, tooling, and platform capabilities
Requirements
- Bachelor’s degree or above; Computer Science, Software Engineering, Network Engineering, Information Security, or related majors preferred
- Five or more years of experience in SRE, DevOps/operations engineering, platform engineering, infrastructure engineering, or related reliability work
- Solid Linux and networking fundamentals; strong familiarity with common distributed architectures, troubleshooting methods, and performance analysis
- Hands-on experience building and operating observability stacks (monitoring, alerting, logging, tracing)
- Automation skills; ability to improve operations and delivery with Python, Go, Shell, or similar
- Practical experience with at least one of CI/CD, configuration management, containers, or cloud infrastructure, with strong change-risk control
- Strong problem decomposition, prioritization, and cross-team influence; comfortable with on-call and a fast-paced environment
- Strong Chinese and English communication and documentation skills; able to work with English technical materials and cross-regional teams
Preferred Experience
- Core SRE: Reliability for large-scale distributed systems, SLO framework design, incident review mechanisms, or proven reduction of MTTR / change failure rate
- Splunk: Splunk architecture and cluster operations, data onboarding and parsing (props/transforms), dashboard/alert/SPL development, access control, and data governance
- High availability and capacity planning for Kubernetes, cloud/virtualization, or middleware (message queues, cache, databases)
- Hands-on use of Prometheus, Grafana, ELK/OpenTelemetry, or similar observability toolchains
- IaC, GitOps, platform engineering, or security/audit logging collaboration experience
- SRE practice in manufacturing, automotive, high-traffic internet, or large enterprise environments
Email:ITRecruiting@tesla.com"">DL-ITRecruiting@tesla.com
app.mokahr.com/su/sAdIC
关于 LearnKu
推荐文章: