Photon Jobs

SPARK Data Onboarding Engineer- NJ

Photon

SPARK Data Onboarding Engineer- NJ

Posted 2 Days Ago

Be an Early Applicant

In-Office or Remote

Hiring Remotely in United States

Senior level

In-Office or Remote

Hiring Remotely in United States

Senior level

Design, develop, and maintain PySpark applications and ETL pipelines to process, transform, and integrate large-scale datasets from SQL, NoSQL, data lakes, and streaming sources. Optimize Spark job performance, implement robust error handling, and collaborate with data analysts, scientists, and architects using orchestration tools like Airflow or Luigi.

The summary above was generated by AI

Job Title: PySpark Data Engineer

Summary:

We are seeking a skilled PySpark Data Engineer to join our team and drive the development of robust data processing and transformation solutions within our data platform. You will be responsible for designing, implementing, and maintaining PySpark-based applications to handle complex data processing tasks, ensure data quality, and integrate with diverse data sources. The ideal candidate possesses strong PySpark development skills, experience with big data technologies, and the ability to work in a fast-paced, data-driven environment.

Key Responsibilities: Data Engineering Development:

Design, develop, and test PySpark-based applications to process, transform, and analyze large-scale datasets from various sources, including relational databases, NoSQL databases, batch files, and real-time data streams.
Implement efficient data transformation and aggregation using PySpark and relevant big data frameworks.
Develop robust error handling and exception management mechanisms to ensure data integrity and system resilience within Spark jobs.
Optimize PySpark jobs for performance, including partitioning, caching, and tuning of Spark configurations.

Data Analysis and Transformation:

Collaborate with data analysts, data scientists, and data architects to understand data processing requirements and deliver high-quality data solutions.
Analyze and interpret data structures, formats, and relationships to implement effective data transformations using PySpark.
Work with distributed datasets in Spark, ensuring optimal performance for large-scale data processing and analytics.

Data Integration and ETL:

Design and implement ETL (Extract, Transform, Load) processes to ingest and integrate data from various sources, ensuring consistency, accuracy, and performance.
Integrate PySpark applications with data sources such as SQL databases, NoSQL databases, data lakes, and streaming platforms

Qualifications and Skills:

Bachelor's degree in Computer Science, Information Technology, or a related field.
5+ years of hands-on experience in big data development, preferably with exposure to data-intensive applications.
Strong understanding of data processing principles, techniques, and best practices in a big data environment.
Proficiency in PySpark, Apache Spark, and related big data technologies for data processing, analysis, and integration.
Experience with ETL development and data pipeline orchestration tools (e.g., Apache Airflow, Luigi).
Strong analytical and problem-solving skills, with the ability to translate business requirements into technical solutions.
Excellent communication and collaboration skills to work effectively with data analysts, data architects, and other team members.

New York, United States

Similar Jobs

Optum

RN Case Manager - Remote California

3 Hours Ago

In-Office or Remote

29-52 Hourly

Junior

29-52 Hourly

Junior

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics

Provide telephonic nursing case management including assessment, care planning, coordination of discharge and follow-up, high-risk patient monitoring, motivational interviewing, documentation in the EHR, and collaboration with physicians and care teams to reduce readmissions and optimize outcomes.

Top Skills: Care Management DashboardsElectronic Health Record

Optum

Remote RN Supervisor I - Cancer Services - Hematology/Oncology - Kelsey Seybold Clinics: Main Campus

3 Hours Ago

In-Office or Remote

73K-130K Annually

Mid level

73K-130K Annually

Mid level

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics

Oversees daily clinical operations for Hematology/Oncology, Radiation Oncology, and Infusion Center services. Supervises staff, ensures competent compassionate patient care, supports accreditation, coordinates multidisciplinary teams, and may cross-cover centers or travel as needed.

Top Skills: AriaBeaconEpicMS Office

Optum

Machine Learning Engineer

3 Hours Ago

In-Office or Remote

165K-282K Annually

Senior level

165K-282K Annually

Senior level

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics

Customer-facing, hands-on AI builder who shapes strategic pursuits, designs end-to-end AI/GenAI architectures, prototypes demos and proofs-of-value, mentors solution leads, supports RFPs, and partners with engineering for implementation and delivery rotations.

Top Skills: Ai AgentsAi FrameworksGenaiLlmsPythonRag

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
Key Industries: Artificial intelligence, Fintech
Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Photon

SPARK Data Onboarding Engineer- NJ

Photon New York, New York, USA Office

Similar Jobs

RN Case Manager - Remote California

Remote RN Supervisor I - Cancer Services - Hematology/Oncology - Kelsey Seybold Clinics: Main Campus

Machine Learning Engineer

What you need to know about the NYC Tech Scene

Key Facts About NYC Tech