Wynd Labs Logo

Wynd Labs

Web Scraping Specialist

Reposted 19 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Lead development and optimization of web scraping pipelines to extract, clean, and store large-scale web data. Handle dynamic content, pagination, distributed scraping, database design with NoSQL, deploy jobs to cloud, and monitor systems for reliability and data quality. Support ML-based data cleaning and categorization.
The summary above was generated by AI

Who We Are:

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.

We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.

We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.

The Role.

We are seeking a Web Scraping Specialist who is proficient and brings significant experience in data extraction and web scraping techniques. You will join a small, specialized team and lead efforts to gather and analyze data, optimize scraping processes, and support our vision for a future where Grass plays a crucial role in transforming internet data accessibility.

Please note: This role requires a work schedule that overlaps sufficiently with EST business hours (min. 3-4 hours) to collaborate effectively with the team.

Who You Are.

  • Demonstrated ability to extract data from complex websites with minimal supervision, with a portfolio or examples of past projects.

  • Proficiency in languages such as Python or JavaScript, with strong skills in libraries and frameworks like BeautifulSoup, Scrapy, or Selenium.

  • Knowledge of asynchronous programming, multithreading, and distributed scraping.

  • In-depth knowledge of HTML, CSS, JavaScript, and the Document Object Model (DOM).

  • Experience with NoSQL databases (MongoDB, Cassandra), capable of designing efficient storage solutions and managing data integrity.

  • Ability to apply machine learning algorithms for data cleaning, categorization, or predictive analysis adds significant value.

  • Experience with cloud services (AWS, Google Cloud, Azure) for deploying and managing scraping jobs at scale.

  • Active participation in open-source projects related to web scraping, data processing, or similar fields.

What You'll Be Doing.

  • Write, test, and refine code that extracts data from various online sources, ensuring reliability and efficiency.

  • Perform data retrieval tasks, handling complexities such as pagination and dynamic content loaded with AJAX.

  • Clean and format extracted data, ensuring it meets quality standards for further analysis or processing.

  • Database management: Store and manage the scraped data in appropriate databases, optimizing for access speed and data integrity.

  • Regularly monitor the scraping processes, identify and resolve any issues to maintain continuous data flow.

Why Work With Us:

  • Opportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.

  • Culture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better.
    We prioritize low ego and high output. This is a fully remote team.

  • Compensation. You’ll receive a competitive salary, benefits and equity package.

Similar Jobs

A Minute Ago
Remote
United States
85K-95K Annually
Senior level
85K-95K Annually
Senior level
Artificial Intelligence • Information Technology • Professional Services • Software • Analytics • Generative AI • Big Data Analytics
Lead execution and optimization of multi-channel ABM programs (1:1, 1:few, 1:many), align campaigns with sales and field marketing, manage paid media, maintain account segmentation and intent data, monitor account-level engagement and pipeline influence, and continuously test and optimize using AI tools. Mentor one team member and coordinate creative, analytics, and execution across internal teams and partners.
Top Skills: Abm PlatformsAnthropic ClaudeCRMDashboards/BiDisplay Advertising PlatformsIntent Data PlatformsMarketing Automation ToolsOpenai ChatgptPaid Search PlatformsPaid Social Platforms
4 Minutes Ago
Easy Apply
Remote or Hybrid
Easy Apply
125K-172K Annually
Senior level
125K-172K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Machine Learning • Software • Big Data Analytics • Automation
Lead executive-level customer relationships to drive adoption of PagerDuty products, deliver measurable business value, and mitigate risk. Create and execute adoption and ROI plans, guide process/people change, coordinate cross-functional post-sales teams, and facilitate reviews, trainings, and strategic engagements. Forecast renewals and expansion while representing customer needs to product and sales teams. Travel up to 25% for in-person meetings.
Top Skills: Ai/MlAutomationCloudDevOpsIt MonitoringPagerduty
7 Minutes Ago
Remote or Hybrid
147K-259K Annually
Expert/Leader
147K-259K Annually
Expert/Leader
Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Develop and execute UVM/SystemVerilog-based verification for AR display integrated circuits. Build assertion-based testbenches, create and run verification plans with functional and code coverage, use Siemens Questa for simulation and debug, and automate verification flows while collaborating with digital, analog, software, and verification teams.
Top Skills: AmbaAsicAssertion-Based TestbenchesCode CoverageEmbedded MicrocontrollerEmulationFunctional CoverageI2CLinuxMakeMipiPerlPythonRtlShellSiemens QuestaSpiSystemverilogTclUvmVerilog

What you need to know about the NYC Tech Scene

As the undisputed financial capital of the world, New York City is an epicenter of startup funding activity. The city has a thriving fintech scene and is a major player in verticals ranging from AI to biotech, cybersecurity and digital media. It also has universities like NYU, Columbia and Cornell Tech attracting students and researchers from across the globe, providing the ecosystem with a constant influx of world-class talent. And its East Coast location and three international airports make it a perfect spot for European companies establishing a foothold in the United States.

Key Facts About NYC Tech

  • Number of Tech Workers: 549,200; 6% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Capgemini, Bloomberg, IBM, Spotify
  • Key Industries: Artificial intelligence, Fintech
  • Funding Landscape: $25.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Greycroft, Thrive Capital, Union Square Ventures, FirstMark Capital, Tiger Global Management, Tribeca Venture Partners, Insight Partners, Two Sigma Ventures
  • Research Centers and Universities: Columbia University, New York University, Fordham University, CUNY, AI Now Institute, Flatiron Institute, C.N. Yang Institute for Theoretical Physics, NASA Space Radiation Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account