Field notes from the world of data extraction.

Articles, interviews and analysis on how data is gathered, used and fought over — written by the people closest to it.

How to use XPath to extract web data
Data quality

Spidermon: Zyte's secret to data quality

Unveil Spidermon's role in our data quality assurance. Elevate confidence in the reliability of your web-scraped data.

Ian Kerins5 min read
How to use XPath to extract web data
Open-source

Meet Spidermon: Our battle tested spider monitoring library

Learn how to easily monitor scrapy spiders & validate data with Spidermon! Developed by Zyte Spidermon is now available as an open-source library.

Renne Rocha6 min read
Proxy management: In-house or off-the-shelf proxy solutions?
Proxies

Proxy management: In-house or off-the-shelf proxy solutions?

This article helps you decide whether you should handle your proxy management in house or by using third party proxy solutions like scraping APIs.

Ian Kerins7 min read
Backconnect proxies explained: How to use them in a scraping project?
Proxies

A sneak peek inside Zyte Smart Proxy Manager

Take a look behind the scenes inside Smart Proxy Manager , the world's smartest web scraping proxy network.

Pablo Hoffman6 min read
Backconnect proxies explained: How to use them in a scraping project?
Product Update

The Zyte Smart Proxy Manager Story: Enhancing Scraping Efficiency

Discover Smart Proxy Manager (Crawlera), the world's smartest proxy network tailored for web scraping, eliminating proxy management hassles.

Pablo Hoffman6 min read
The rise of web data in hedge fund decision making
Web data application

The rise of web data in hedge fund decision making

Learn about the rise of web data in hedge fund decision making and the importance of data quality in this blog post.

Ian Kerins8 min read
Witness the power of web scraped product data for investors
Use case

Witness the power of web scraped product data for investors

Investors understand the importance of high-quality information. Learn how Eagle Alpha predicted GoPro's quarterly revenues using web scraped alternative data.

Ian Kerins6 min read
Autoscraping Casts A Wider Net
Developer interest

Looking Back at 2018: A Year of Growth and Challenges

Looking Back at 2018 - Join us in reminiscing about the highlights and achievements of the year 2018.

Ian Kerins4 min read
Blog Comments API Beta Release: Engaging with Readers
Leadership

GDPR and Web Scraping: IIAP Europe Data Protection Congress

GDPR & Web Scraping: IIAP Europe Data Protection Congress - Learn about GDPR's impact on web scraping and data protection from industry experts at IIAP Europe Data Protection Congress.

Sanaea Daruwalla4 min read
EuroPython 2015: Uniting Pythonistas in Europe
Developer interest

Shubber Gettogether 2018: Gathering Innovators and Thinkers

Shubber Gettogether 2018 - Relive the excitement of the Shubber Gettogether in 2018, where experts gathered to discuss cutting-edge web scraping techniques.

Ian Kerins3 min read
A Career in Remote Working: Embracing Flexibility and Freedom
Data quality

Data Quality Assurance For Enterprise Web Scraping

Data quality is essential to the success of any web scraping project, especially when scraping the web at scale.

Ian Kerins7 min read
Google Summer of Code 2015: Empowering Open Source Projects
Developer interest

What I Learned As A Google Summer Of Code Student At Zyte

Google Summer of Code (GSoC) was such a great experience for students like me. I learned so much about open source communities as well as contributing to

Chau Tung Lam Nguyen Bhatt4 min read