January 2020 Summaries
3 posts from Bright Data
Filter
Month:
Year:
Post Summaries
Back to Blog
Web scraping, or data harvesting, is a technique used to extract various types of data, from product information to public records, and can be accomplished using different tools that may or may not involve proxies. While proxies offer benefits such as reduced risk of being blocked and faster data collection for large-scale operations, small-scale data extraction can often be performed without them. Using methods like slowing down scraping speed, hiding IP addresses with tools like Tor or VPNs, rotating user agents, and employing headless browsers, individuals can attempt to gather data while minimizing detection. However, these methods have limitations in terms of speed and reliability, especially when dealing with large volumes of data. For more efficient and extensive data collection, using proxies is recommended as it allows for scalable access without restrictions, which is crucial for serious data mining efforts.
Jan 22, 2020
1,076 words in the original blog post.
Encountering erroneous proxy messages is a common issue in online data management and web scraping, with error codes serving as vital indicators for diagnosing and resolving data delivery problems. The article provides an overview of various HTTP proxy error codes, categorizing them into client-side and server-side errors, and explaining their meanings and typical conditions under which they arise. It covers 3xx codes like 301 and 302 for redirects, 4xx codes such as 400 for bad requests and 403 for forbidden access, and 5xx codes like 500 for server errors. Each error code is accompanied by practical advice on troubleshooting and solutions, such as handling redirects, verifying request syntax, and implementing retry strategies to reduce the incidence of errors. Additionally, the article emphasizes the importance of understanding proxy settings and consulting documentation to effectively manage HTTP errors during web scraping.
Jan 09, 2020
2,682 words in the original blog post.
Bright Data has significantly enhanced its network architecture, achieving a 25% increase in speed, making its Datacenter IPs the fastest on the market. A proxy, which is essentially an IP address, plays a crucial role in network speed, influenced by factors like journey length, architecture, bandwidth, and CPU usage. To optimize performance, Bright Data has minimized the journey length by ensuring requests remain within the same geolocation, reduced unnecessary hops, and prevented server overloads by distributing requests across multiple load-bearing servers. Additionally, they have ensured ample bandwidth and CPU resources to avoid bottlenecks. These improvements cater to the needs of customers who require fast proxies for various use cases, reaffirming Bright Data's commitment to providing the fastest proxies available. Aviv Besinsky, a lead product manager at Bright Data, has been instrumental in advancing data collection technology and shares his expertise in the field.
Jan 05, 2020
721 words in the original blog post.