D

Databricks

Data Engineer (Mock Interview)

Interview Date

May 16, 2025 (implied by video upload date)

Result

Not Specified

Difficulty

Rounds

Not explicitly defined as "rounds" but covers various technical discussions and coding tasks.

Drive Type

Full-Time

Topics asked

Challenges with DatabricksAzure SynapseAzure Data Factory (ADF)SQLPythonPySparkSpark architectureSpark optimization (repartition vs coalesce)CTE vs SubqueriesSQL query optimizationAI applications in projectsPython package creationfetching data from Gold layer (Python APIframework)SQL (third highest salary)Python (count wordsreverse string)Slowly Changing Dimensions (SCD)Schema EvolutionScenario-Based Questions (Azure)Microsoft Fabric vs other data engineering servicesCaching (Redis).

Detailed experience

Role: Data Engineer (Mock Interview)

College: Not Specified (Masters in 2022)

Interview Date: May 16, 2025 (implied by video upload date)

Interview Type: Full-Time

Result: Not Specified

Difficulty: Not Specified

Rounds: Not explicitly defined as "rounds" but covers various technical discussions and coding tasks.

Topics Asked: Challenges with Databricks, Azure Synapse, Azure Data Factory (ADF), SQL, Python, PySpark, Spark architecture, Spark optimization (repartition vs coalesce), CTE vs Subqueries, SQL query optimization, AI applications in projects, Python package creation, fetching data from Gold layer (Python API/framework), SQL (third highest salary), Python (count words, reverse string), Slowly Changing Dimensions (SCD), Schema Evolution, Scenario-Based Questions (Azure), Microsoft Fabric vs other data engineering services, Caching (Redis).

Experience:

This is a mock interview experience for a Data Engineer role, reflecting real-world scenarios and commonly asked questions. The candidate had around 2 years and 9 months of experience, primarily in data engineering using Azure services (Azure Synapse, Azure Databricks, Azure Data Factory) along with Python, SQL, and PySpark.

The interview covered a broad range of data engineering topics:

  • **Project Discussions:** Questions about challenges faced while working on Databricks, and detailed discussions on Azure Synapse, including its use as a data warehouse and data transfer from on-premises.
  • **SQL and Python:** How these languages were used in projects, differences between CTEs and subqueries, and SQL query optimization techniques. Specific coding questions included finding the third highest salary in SQL and Python tasks like counting words in a string and reversing a string.
  • **Spark and Data Warehousing:** Spark architecture, optimizing Spark jobs (e.g., repartition() vs. coalesce()), and services used in data engineering projects. Discussions also touched upon fetching data from the Gold layer using Python APIs/frameworks.
  • **Data Modeling and ETL:** Questions on Slowly Changing Dimensions (SCD) and Schema Evolution.
  • **Cloud and Databricks Specifics:** Scenario-based questions related to Azure, the difference between Microsoft Fabric and other data engineering services, and general knowledge of caching mechanisms like Redis. The candidate also discussed using Databricks notebooks for modeling (star schema) within a medallion architecture.

The candidate had some exposure to Databricks for transformations from silver to gold layers within a medallion architecture but noted not having extensive hands-on experience with all Databricks features.

Posted on - 12 Nov 2025