Build and maintain production-grade batch and streaming data pipelines in Databricks using PySpark and SQL. Develop ingestion frameworks, data models, APIs, data quality checks, monitoring, and Medallion Architecture across bronze, silver, and gold layers. Integrate diverse data sources, optimize Spark workloads, document data lineage and architecture, and collaborate with engineering and federal stakeholders. An active Secret clearance or higher is required.
This is a remote position.
Location: Remote – United States
Security Clearance: Active Secret or higher REQUIRED
Our client, a growing technology services organization supporting federal government programs, is seeking an experienced Databricks Data Engineer to support a federal technology initiative.
This is a remote opportunity available to U.S. citizens residing in the United States. Candidates must currently hold an ACTIVE Secret security clearance or higher. Candidates without an active Secret-level clearance cannot be considered.
The ideal candidate will bring strong hands-on data engineering experience within Databricks, including development of production-grade batch and streaming pipelines, PySpark and SQL transformations, data modeling, ingestion frameworks, and modern data architecture practices.
- Design, build, and maintain batch and streaming data pipelines using PySpark, SQL, Databricks Workflows, and Delta Live Tables.
- Implement Medallion Architecture across bronze, silver, and gold data layers to support data quality, transformation, and consumption.
- Develop scalable ingestion frameworks for structured, semi-structured, and unstructured data.
- Integrate data from files, databases, APIs, and streaming sources such as Kafka, Kinesis, and Databricks Auto Loader.
- Design dimensional and domain-specific data models supporting analytics and downstream applications.
- Build and consume APIs for integration with downstream systems.
- Optimize Spark workloads for performance and cost through partitioning, caching, cluster sizing, and related techniques.
- Develop and implement data quality checks, validation processes, and pipeline monitoring.
- Maintain documentation covering data flows, lineage, architecture, and integration points.
- Collaborate with engineering, platform, architecture, and federal program stakeholders throughout the development lifecycle.
Requirements
- Active Secret security clearance or higher is required.
- 5+ years of professional data engineering experience.
- 2+ years of hands-on Databricks experience.
- Strong production-level experience with PySpark and SQL.
- Experience developing both batch and streaming data pipelines.
- Strong understanding of data modeling and modern data architecture principles.
- Proficiency with Python.
- Experience working within Git-based development and CI/CD environments.
- Familiarity with Databricks Unity Catalog, including catalogs, schemas, and permissions from a data engineering perspective.
- Experience integrating data from multiple source types, including databases, APIs, files, and streaming platforms.
- Experience with Databricks Delta Live Tables.
- Experience with Databricks Workflows.
- Experience with Kafka, Kinesis, or similar streaming technologies.
- Experience working within federal, regulated, or security-sensitive environments.
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.
Similar Jobs
Big Data • Cloud • Information Technology • Analytics • Business Intelligence • Consulting • Data Privacy
Designs and implements scalable data, analytics, and AI solutions for clients. Builds and optimizes data models and ETL/ELT pipelines, performs testing and validation, translates business requirements into technical solutions, documents architectures, manages workstreams, supports client training, and contributes to governance, risk tracking, and solution design.
Top Skills:
Amazon RedshiftAzure SynapseBigQueryDatabricksEltETLPythonSnowflakeSQL
Financial Services
Build and operate scalable Databricks-on-AWS data pipelines using PySpark, Delta Lake, and lakehouse patterns. Optimize performance, implement data quality, monitoring, alerting, and automated remediation, and deliver curated datasets for BI and analytics partners. Collaborate with stakeholders on architecture and design while applying secure software engineering, CI/CD, agile, and operational stability practices. The role also uses AI-assisted development tools and supports workforce data analytics.
Top Skills:
AlteryxAmazon AthenaAmazon EmrAmazon S3Apache AirflowApache IcebergSparkAutosysAWSAws CloudwatchAws GlueAws LambdaBitbucketClaudeDatabricksDatabricks WorkflowsDelta LakeDelta Live TablesGitGithub CopilotJavaJenkinsOracleParquetPysparkPythonScalaSigmaSpinnakerSQLTableau
Software
Designs, builds, and operates scalable batch and streaming data pipelines on Databricks for federal missions. Responsibilities include developing Spark and Delta Lake solutions, managing clusters and workflows, implementing Unity Catalog governance and security, optimizing ETL/ELT processes, integrating CI/CD, supporting machine learning and advanced analytics, monitoring data quality, and collaborating with technical teams and stakeholders.
Top Skills:
Amazon EmrSparkAWSAzureCi/CdDatabricksDelta LakeGgplot2GitGCPHadoopHiveKafkaMlflowNoSQLPlotlyPysparkPythonSeabornSpark SqlSpark Structured StreamingSQLUnity Catalog
What you need to know about the NYC Tech Scene
As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.
Key Facts About NYC Tech
- Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
- Key Industries: Artificial intelligence, Fintech
- Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
- Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory


