INTERVIEW QUESTIONS · DATA ENGINEER

Data engineer interview questions and how to answer them

Questions on the skills open data engineer listings mention and on the core work of the role. For each: what the interviewer is checking, and a STAR answer outline with placeholders for your own example.

See open data engineer rolesSee Interview Studio

Questions on the skills these listings mention

Based on 104 open data engineer listings on ROLIVA as of . Each question below is about a term from ROLIVA’s skills vocabulary that appears in at least two of these listings, most-mentioned first.

01 Tell me about something you built or automated in Python that other people relied on. What would you do differently now?

Why this question: Python is mentioned in 85 of 104 listings.

What the interviewer is checking. Whether you write Python that works beyond your own laptop: structure, error handling, tests or checks, and whether you can judge your earlier work honestly.

Answer outline (STAR)

  1. Situation. [The task or problem, and who depended on the script, notebook or service].
  2. Task. [What it had to do and any constraint, e.g. data size, run frequency or a deadline].
  3. Action. [The approach and libraries you used, how you handled failures or bad input, and how you tested it].
  4. Result. [What it replaced or made possible, and the one thing you would change today]. If there is no number you can support, say what changed as a result or what you learned.

02 Walk me through a SQL query you wrote that answered a real question. How did you know the result was right?

Why this question: SQL is mentioned in 84 of 104 listings.

What the interviewer is checking. Whether you can turn a question into joins, filters and aggregations, and whether you check your own output instead of trusting the first number the query returns.

Answer outline (STAR)

  1. Situation. [The question someone needed answered and the tables or data sources involved, e.g. orders joined to customers].
  2. Task. [What you had to produce, for whom, and by when].
  3. Action. [How you built the query: the joins, filters and grouping you chose, and how you checked it, e.g. row counts, a hand-checked sample or a comparison with a known total].
  4. Result. [What the answer showed and what was decided because of it]. If there is no number you can support, say what changed as a result or what you learned.

03 Describe something you built or ran on AWS. Which services did you choose, and what went wrong at some point?

Why this question: AWS is mentioned in 67 of 104 listings.

What the interviewer is checking. Whether you understand the trade-offs behind the services you used, including security, cost and failure modes, rather than just naming them.

Answer outline (STAR)

  1. Situation. [The system and what it did, e.g. an API, a data pipeline or a batch job].
  2. Task. [Your part in building or operating it].
  3. Action. [The services you used and why, how you handled permissions and cost, and how you dealt with the failure].
  4. Result. [How the system ran afterwards and what you changed to prevent a repeat]. If there is no number you can support, say what changed as a result or what you learned.

04 Walk me through a Spark job you built or tuned. What made it slow or expensive, and how did you fix it?

Why this question: Spark is mentioned in 53 of 104 listings.

What the interviewer is checking. Whether you understand distributed processing (partitions, shuffles, skew, memory) and can tune by evidence rather than by guesswork.

Answer outline (STAR)

  1. Situation. [The job, its input size and what consumed its output].
  2. Task. [The problem, e.g. a job missing its window, failing on memory or costing too much].
  3. Action. [How you diagnosed it, e.g. the Spark UI or stage metrics, and what you changed, e.g. partitioning, a broadcast join or handling skew].
  4. Result. [Run time, reliability or cost before and after]. If there is no number you can support, say what changed as a result or what you learned.

05 Walk me through work you did on Azure. How did you handle identity, access and environments?

Why this question: Azure is mentioned in 52 of 104 listings.

What the interviewer is checking. Whether you can work safely in a cloud platform: least-privilege access, separate environments, and repeatable deployment instead of manual changes.

Answer outline (STAR)

  1. Situation. [The system or migration and the Azure services involved].
  2. Task. [What you were responsible for and the constraint, e.g. a compliance requirement].
  3. Action. [How you set up access, environments and deployment, and any problem you solved along the way].
  4. Result. [The result, e.g. a successful migration or fewer manual changes]. If there is no number you can support, say what changed as a result or what you learned.

06 Tell me about work you did in Snowflake where performance or cost mattered. What did you change and why?

Why this question: Snowflake is mentioned in 51 of 104 listings.

What the interviewer is checking. Whether you understand how warehouse size, query design and storage choices affect speed and cost, and whether you measure before and after.

Answer outline (STAR)

  1. Situation. [The workload, e.g. a slow transformation or an expensive scheduled query].
  2. Task. [What you were asked to improve and the constraint, e.g. a reporting deadline or a budget].
  3. Action. [What you investigated, e.g. query profile or warehouse usage, and the change you made].
  4. Result. [The effect on run time or cost, measured the same way before and after]. If there is no number you can support, say what changed as a result or what you learned.

07 Tell me about a pipeline you orchestrated with Airflow. What happened when a task failed in the middle of the night?

Why this question: Airflow is mentioned in 50 of 104 listings.

What the interviewer is checking. Whether you design pipelines that fail safely and can be re-run: idempotent tasks, sensible retries, alerting and clear dependencies.

Answer outline (STAR)

  1. Situation. [The pipeline and who relied on its output].
  2. Task. [The failure or the reliability goal].
  3. Action. [How you designed or changed the DAG, e.g. idempotent tasks, retries, alerts or backfills, and how you recovered].
  4. Result. [How reliability changed and what downstream users noticed]. If there is no number you can support, say what changed as a result or what you learned.

08 Tell me about a machine learning model you worked on that was used for a real decision. How did you know it was good enough?

Why this question: Machine learning is mentioned in 44 of 104 listings.

What the interviewer is checking. Whether you frame the problem, choose a baseline and metrics that match the decision, and check for leakage and drift rather than reporting a single score.

Answer outline (STAR)

  1. Situation. [The decision the model supported and the data available].
  2. Task. [Your part: framing, features, modelling, evaluation or deployment].
  3. Action. [The baseline, the model, how you evaluated it, e.g. a held-out set or an online test, and the risks you checked].
  4. Result. [How it was used and how it performed after release]. If there is no number you can support, say what changed as a result or what you learned.

Questions about the core work of a data engineer

These follow from the work the role title describes, not from a count of listings.

09 Tell me about a data pipeline you designed. What decisions did you make about how it handles failure?

What the interviewer is checking. Design thinking for reliability: idempotent loads, retries, backfills, data quality checks and alerting, not just moving data from A to B.

Answer outline (STAR)

  1. Situation. Describe the pipeline: [the sources, the destination, the volume and who depended on it].
  2. Task. State what you owned: [e.g. the design, build and on-call support].
  3. Action. Explain your design choices [e.g. batch or streaming and why; how reruns avoid duplicates; the data quality checks; alerts and who receives them].
  4. Result. Give a result you can support [e.g. reliability over a period, faster recovery]. If it failed in an unexpected way, say what you changed.

10 Describe a time a broken upstream change corrupted data your users relied on. What did you do?

What the interviewer is checking. Incident handling, communication with data consumers, and the contracts or checks you put in place to stop it happening again.

Answer outline (STAR)

  1. Situation. Describe the incident: [the upstream change, the data affected and how it was discovered].
  2. Task. Say what you were responsible for: [e.g. stopping the bad data, repair, informing users].
  3. Action. Explain your actions [how you contained it, how you repaired or backfilled the data, who you informed and how, and the prevention you added, e.g. schema checks or an agreement with the source team].
  4. Result. Share the result you can support [e.g. time to repair, no repeat]. If users had already used the bad data, say how you helped them correct it.

11 How have you improved the cost or performance of a slow or expensive data workload?

What the interviewer is checking. Whether you measure before optimising, understand where cost and time go, and balance performance against maintainability.

Answer outline (STAR)

  1. Situation. Describe the workload: [the job or query, how slow or expensive it was, and who noticed].
  2. Task. State your goal: [e.g. finish before a reporting deadline or reduce cost].
  3. Action. Explain how you found the cause [the profiling or monitoring you used] and what you changed [e.g. partitioning, incremental processing, query rewrites].
  4. Result. Give the result you can support [e.g. runtime or cost before and after]. If the gain was small, say what you learned about where the bottleneck really was.

12 Tell me about a time analysts or scientists disagreed with how you modelled a dataset.

What the interviewer is checking. Collaboration with data consumers, understanding of modelling trade-offs, and willingness to change your design based on how data is used.

Answer outline (STAR)

  1. Situation. Describe the dataset and the disagreement: [your model, and what the users wanted instead].
  2. Task. State what you needed to reach: [e.g. one agreed model that served several teams].
  3. Action. Explain what you did [how you understood their queries and needs, the options you proposed, and how you reached a decision].
  4. Result. Share the result you can support [what was built and how it was used]. If you kept your design, say how you explained it and what you documented.

Using the outlines

How these questions are chosen

Each open data engineer listing’s description and stated skills are checked, on whole words, against the fixed skills vocabulary ROLIVA uses on its job pages, resume examples and monthly skills reports. A listing counts once per term. A mention is not a requirement, and the counts change as employers open and close roles. The questions and outlines are preparation prompts written for this role. They are not a record of what any employer has asked, and no employer’s process is described here.

Further reading on the STAR method and interview preparation (checked 5 October 2026): National Careers Service: interview advice · Job Bank: prepare for an interview

Related

Prepare for a specific job.

Interview Studio prepares questions and STAR outlines for a real job from its description and the experience you have confirmed. No generative AI: when a detail is missing, it asks you instead of inventing one.

Create my free workspace