Tóm lược
- Yêu cầu kỹ thuật:
- Data Modeling ,
- ETL ,
- ELT ,
- Java ,
- PostgreSQL ,
- Python ,
- OLAP ,
- MS SQL ,
- Distributed Systems ,
- REST API ,
- Elasticsearch ,
- Docker ,
- Apache Spark ,
- Redis ,
- Architecture ,
- Observability ,
- AWS Kinesis ,
- AWS Redshift ,
- Grafana ,
- Amazon S3 ,
- Amazon RDS ,
- Apache Kafka ,
- Stream processing ,
- AWS ,
- Amazon SQS ,
- Kotlin ,
- IoT ,
- Kubernetes ,
- OLTP ,
- DataDog ,
- Orchestrator ,
- Apache Airflow ,
- Dagster ,
- Prefect ,
- Terraform ,
- Apache Flink ,
- DBT ,
- ClickHouse ,
- Druid ,
- LLM ,
- Vector ,
- Amazon Aurora ,
- Google BigQuery ,
- Amazon EKS ,
- CI/CD ,
- Prometheus ,
- Elastic Stack ,
- ELK ,
- Cursor ,
- Claude ,
- Codex ,
- Gemini
Mô tả công việc
Tóm tắt công việc
Build like a founder. Ship like Spartan.
Our clients are startups in a race: an MVP in 30 days, and billions of data points the year after. We're looking for a Data Engineer to build the platform that carries it.
We don't act like vendors. We embed. Our teams sit inside startups backed by YC, Pear VC, BlingCap, a16z, and Techstars, working beside founders and engineers from Stanford, MIT, Yale, Google, Amazon, and Uber. Together we ship those MVPs and scale the systems behind them to millions of users and billions of requests.
This isn't a batch-and-dashboards seat three layers from the client. You'll be in the founder's Slack. Your pipelines run in production. You'll own the data platform behind real-time operations, analytics, and AI, from ingestion all the way to serving.
Data is the core of the job. Backend and infra come with it, because owning the path to production means you deploy your own work and build the services that hand data to the product. You don't need to arrive as an expert in all three. You do need to be willing to learn the other two.
70% of our portfolio is LLM-powered, and every one of those features runs on data someone had to make trustworthy first. That's your work. Cursor, Claude, and Codex are default tools here, not experiments. You'll use AI every day, and you'll be expected to use it well.
Who Thrives at Spartan
Not every environment fits everyone, so here's an honest picture of the teammate we're looking for.
You'll thrive here if:
- You do whatever it takes to win. A schema redesign today, a backfill script and a founder demo tomorrow.
- You take on the hardest part of the project and see it through to shipped.
- You chase a wrong number until you find where it broke. Silent bad data bothers you more than a loud outage.
- You'd rather own a problem than a ticket.
Some weeks will stretch you. That's the job. But you'll never carry it alone. A critical client release or a tricky production issue? Someone always has your back. We push each other hard because we care how far each of us can go. We win together. We lose together. Nobody gets left behind.
One last thing. We're not filling seats. We're looking for teammates for the battles ahead. If you read all this and something in you said yes, trust that feeling and apply. We respect people who back themselves. Every Spartan here started with that same yes.
What You'll Do
- Build ingestion for operational, telemetry, and business data: ETL and ELT, incremental loads, CDC, backfills that don't take the system down.
- Model data for how it gets used: event schemas upstream, star schemas and marts downstream.
- Build streaming pipelines that land data in near real time: windowed aggregations, stream joins, late and duplicate events, exactly-once writes to the sink.
- Own data infrastructure end to end: ingestion, processing, storage, serving.
- Own data quality: contracts, schema changes, lineage, and alerts that fire before a founder spots a wrong number.
- Build the data layer behind AI features: feature pipelines, embeddings and vector stores, RAG source data, eval datasets.
- Build the serving layer: APIs and services that expose metrics, aggregates, and feature data to product teams.
- Ship fast with a proven process: async standups, PRDs/RFCs, code reviews, and an AI-native workflow on Cursor, Claude, and Codex.
- Talk directly with US startup founders and product leaders. Your input shapes what gets measured and what gets built.
- Wear more than one hat. Data first, but you'll touch backend, infra, and AI integration when the project needs it.
- Help shape Spartan's engineering culture. We work like founders, not contractors.
Career Growth
- Work directly with founders from YC, Stanford, MIT, Yale, Google, Amazon, Uber, and more.
- Exposure to streaming, real-time systems, AI infrastructure, and cloud at scale.
- Transparent career path and mentorship from engineers who've taken 3 companies to IPO.
- Every project feels like its own mini-startup.
Why Spartan
- 50+ MVPs launched, 3 IPOs, $100M+ raised by the startups we build with.
- 95% engineer retention. People stay because the work is worth staying for.
- You build zero to one. First event ingested to a live data platform in weeks. Most engineers get that chance once in a career. Here it's the job.
- Your work is visible. The pipeline you build this month is the number in a founder's investor deck next month.
- No layers, no bureaucracy. You talk to the decision-maker directly, and decisions land in hours, not weeks.
- Every project is a new market, a new stack, a new problem. One year here teaches what three years elsewhere would.
- You'll never just take tickets. You'll own problems, ship fast, and grow with the founders we partner with.
Benefits & Perks
- 💰 Competitive salary, 100% salary during probation
- 🏝 Unlimited PTO + public holidays
- 🩺 Premium healthcare insurance + annual health check
- ⚽ Sports clubs for every taste: football, badminton, billiards, pickleball
- 🎮 PS5 and Nintendo at the office
- 🏃 Strava yearly subscription
- 🏠 Hybrid working style, flexible hours
- ✈️ Annual company trip + monthly team parties
- 💻 MacBook provided
- 📜 Social insurance per Labor Law
Yêu cầu công việc
- Fluent English. You'll talk directly with US startup founders.
- 3+ years building and running production data pipelines: ingestion, ETL and ELT, incremental loads, backfills, and a plan for when a run fails halfway through.
- Data modeling you can defend: normalized schemas for OLTP, dimensional models for analytics (star schema, fact and dimension tables, slowly changing dimensions), and a reason for when you denormalize instead.
- Strong SQL. You can read a query plan and know why the job got slow.
- You've run pipelines on an orchestrator in production: Airflow, Dagster, Prefect, or similar. Scheduling, dependencies, retries, idempotent reruns.
- Strong coding skills in Python, Java, Kotlin, or similar. You write tested, reviewable code, not one-off scripts.
- Hands-on with streaming data in production: Kafka, Pulsar, Kinesis, or similar. You've dealt with consumer lag, replays, and duplicate events.
- Comfortable in the cloud. You deploy and debug your own pipelines on AWS, and Terraform or a Kubernetes manifest doesn't scare you.
- Strong CS fundamentals: distributed systems, partitioning, and what breaks when the data outgrows one machine.
- You use AI tools like a pro: you review what they write, catch their mistakes, and know when not to use them.
- A builder's mindset. You don't just move data; you build platforms that make data useful.
Nice to Have
- Stream processing in production: Flink, Spark, or Kafka Streams.
- Warehouse and lakehouse work: dbt, Iceberg or Delta, or a real-time OLAP store like ClickHouse or Druid.
- Durable execution for long-running workflows: Temporal or similar.
- IoT or telemetry data at scale: high-volume ingest, late and out-of-order events, device fleets.
- ML or LLM infrastructure: feature stores, vector databases, training or eval data pipelines.
- You've built an internal data platform from scratch.
- Domain experience in rideshare, logistics, mobility, transportation, or another data-heavy vertical.
Tech We Use
Every project brings a different stack, so in a year here you'll touch more of this list than most engineers do in five. What we adopt is decided by our engineers through an internal Tech Radar, not handed down from a slide deck.
- Languages: Kotlin/Java · Python · SQL
- Streaming: Kafka · Pulsar · AWS Kinesis · SQS
- Processing: Flink · Spark
- Orchestration: Airflow · Temporal (durable execution)
- Storage: PostgreSQL · Aurora · Redis · S3
- Warehouse & analytics: Redshift · BigQuery · ClickHouse
- Data serving: REST APIs · event-driven architecture
- Vector search: pgvector · Pinecone · Qdrant · Weaviate
- Cloud: AWS (EKS, S3, RDS) · Kubernetes · Docker · Terraform · CI/CD
- Observability: DataDog · Prometheus · Grafana · ELK Stack
- AI tooling: Cursor · Claude · Codex · Gemini
Ngôn ngữ
-
English
Nói: Fluently - Đọc: Fluently - Viết: Fluently
Yêu cầu kỹ thuật
- Data Modeling
- ETL
- ELT
- Java
- PostgreSQL
- Python
- OLAP
- MS SQL
- Distributed Systems
- REST API
- Elasticsearch
- Docker
- Apache Spark
- Redis
- Architecture
- Observability
- AWS Kinesis
- AWS Redshift
- Grafana
- Amazon S3
- Amazon RDS
- Apache Kafka
- Stream processing
- AWS
- Amazon SQS
- Kotlin
- IoT
- Kubernetes
- OLTP
- DataDog
- Orchestrator
- Apache Airflow
- Dagster
- Prefect
- Terraform
- Apache Flink
- DBT
- ClickHouse
- Druid
- LLM
- Vector
- Amazon Aurora
- Google BigQuery
- Amazon EKS
- CI/CD
- Prometheus
- Elastic Stack
- ELK
- Cursor
- Claude
- Codex
- Gemini
Thông tin doanh nghiệp
Spartan is the pioneering technology solution company empowering startups on their path to profitability.
We prioritize the growth and success of our top-notch engineers, leveraging the benefits of remote work to deliver exceptional results while optimizing costs.
At Spartan, we offer a comprehensive range of services tailored to meet the diverse needs of startups, established enterprises, and so on. Our core expertise lies in five primary areas:
- Mobile/Web App Development
- Web Development
- AI/ML Development
- Backend development
- Cloud Solutions
With a robust tech stack that spans from infrastructure to backend and frontend development, we possess the necessary tools and expertise to tackle complex projects. Furthermore, we actively integrate emerging technologies such as AI integration, IoT, and real-time streaming data into our solutions. These technological advancements enable us to deliver exceptional results and stay at the forefront of innovation.