[{"data":1,"prerenderedAt":40},["ShallowReactive",2],{"article":3},{"id":4,"category":5,"slug":6,"title":7,"image":8,"page_image":9,"published_at":10,"updated_at":10,"meta_title":11,"meta_description":12,"meta_keywords":13,"content":14,"translations":15,"tags":36,"faqs":39},150,"blog","gathering-high-quality-web-data-at-scale-5-tweaks-to-apply","Gathering high-quality web data at scale: 5 tweaks to apply","https://blog.dexodata.com/storage/uploads/previews/21-2-s-trusted-proxy-website-gathering-high-quality-web-data-at-scale-5-tweaks-to-apply-cover-4caec7dc-43fe-450e-801a-3bd3b995c628.webp","https://blog.dexodata.com/storage/uploads/covers/21-2-b-trusted-proxy-website-gathering-high-quality-web-data-at-scale-5-tweaks-to-apply-cover-71633833-80fa-440e-9314-29f32310d546.webp","2025/11/21","How to collect high-quality web data at scale with geo targeted proxies","5 steps to high-quality web data harvesting: scraping tools, AI models, and dynamic geo targeted proxies from the Dexodata ecosystem for ethical scraping.","buy residential and mobile proxies, buy residential ip, geo targeted proxies","\u003C!DOCTYPE html PUBLIC \"-//W3C//DTD HTML 4.0 Transitional//EN\" \"http://www.w3.org/TR/REC-html40/loose.dtd\">\n\u003C?xml encoding=\"utf-8\"?>\u003Chtml>\u003Cbody>\u003Cp>\u003Cem>\u003Cstrong>Contents of article:\u003C/strong>\u003C/em>\u003C/p>\r\n\u003Cul>\r\n\u003Cli>\u003Ca href=\"#anchor1\">5 steps to collect high-quality data at scale\u003C/a>\u003C/li>\r\n\u003C/ul>\r\n\u003Cul>\r\n\u003Cli style=\"list-style-type: none;\">\r\n\u003Cul>\r\n\u003Cli>\u003Ca href=\"#anchor2\">What is high-quality data\u003C/a>\u003C/li>\r\n\u003Cli>\u003Ca href=\"#anchor3\">1. Choosing a framework and libraries\u003C/a>\u003C/li>\r\n\u003Cli>\u003Ca href=\"#anchor4\">2. Crawling sites\u003C/a>\u003C/li>\r\n\u003Cli>\u003Ca href=\"#anchor5\">3. Running dynamic proxies\u003C/a>\u003C/li>\r\n\u003Cli>\u003Ca href=\"#anchor6\">4. Cleaning and preprocessing raw data\u003C/a>\u003C/li>\r\n\u003Cli>\u003Ca href=\"#anchor7\">5. Applying machine learning\u003C/a>\u003C/li>\r\n\u003C/ul>\r\n\u003C/li>\r\n\u003Cli>\u003Ca href=\"#anchor8\">High-quality web data collection and Dexodata\u003C/a>\r\n\u003Cp>Awareness of trending technologies is important as well as market situation intelligence. \u003Ca href=\"https://dexodata.com/en/blog/what-is-web-data-extraction\" target=\"_blank\" rel=\"noopener\">Web data harvesting through geo targeted proxies\u003C/a> is the number one procedure to make considered business decisions in e-commerce, ads verification, developing and promoting products or services.\u003C/p>\r\n\u003Cp>The market of scraping solutions grows along with related IT spheres, and its current value is more than \u003Ca href=\"https://www.researchnester.com/reports/web-scraping-software-market/5041\" target=\"_blank\" rel=\"noopener\">$4 billion, with possibility to grow four times to 2035\u003C/a>. Social media, real estate platforms and healthcare are the main drivers of info extraction tools&rsquo; development. Considering the scale of publicly available online data, analysts buy residential and mobile proxies to obtain internet insights seamlessly and ethically. The Dexodata ecosystem offers to buy residential IP pools suitable for corporate-leveled web info acquisition, due to:\u003C/p>\r\n\u003Cul>\r\n\u003Cli>Strict AML and KYC compliance\u003C/li>\r\n\u003Cli>External addresses rotation\u003C/li>\r\n\u003Cli>HTTP and SOCKS5 compatibility\u003C/li>\r\n\u003Cli>Flexible pricing plans with adjustable geolocation and traffic amounts.\u003C/li>\r\n\u003C/ul>\r\n\u003Cp>Implementing an intermediate infrastructure into scraping software is the first step to gathering high-quality web data at scale. We will clarify other steps below.\u003C/p>\r\n\u003Ch2>\u003Ca name=\"anchor1\">\u003C/a>5 steps to collect high-quality data at scale\u003C/h2>\r\n\u003Cp>The essence of scraping lies in creating an automated algorithm, which detects relevant insights on the internet source, obtains them, and places extracted details into .json, .xml, .csv datasets for further analysis. Geo targeted proxies account for delivering HTTP GET or POST requests to the target page, while automated scripts boost and control this scenario.&nbsp;\u003C/p>\r\n\u003Cp>The main steps to perform high-quality scraping include:\u003C/p>\r\n\u003Col>\r\n\u003Cli>Choosing a framework and libraries\u003C/li>\r\n\u003Cli>Crawling sites\u003C/li>\r\n\u003Cli>Running dynamic proxies\u003C/li>\r\n\u003Cli>Cleaning and preprocessing raw data\u003C/li>\r\n\u003Cli>Applying machine learning.\u003C/li>\r\n\u003C/ol>\r\n\u003Cp>Each phase considers dozens of factors &mdash; scale, geographical determinacy, target sources&rsquo; number, API availability, JS-based page structure, CPU performance, project&rsquo;s budget, and more. IT engineers ponder whether to buy residential IPs, mobile or datacenter, to \u003Ca href=\"https://dexodata.com/en/blog/ipv4-or-ipv6-what-to-prefer-for-obtaining-web-data\" target=\"_blank\" rel=\"noopener\">choose IPv4 or IPv6 addresses\u003C/a>. These features are among those which determine the quality of gathered information.\u003C/p>\r\n\u003Ch3>\u003Ca name=\"anchor2\">\u003C/a>What is high-quality web data?\u003C/h3>\r\n\u003Cp>There are key metrics showing to what extent the obtained info suits defined objectives. The higher are values, the more accurate and actual will be business decisions made on its basis. These parameters are:\u003C/p>\r\n\u003Cul>\r\n\u003Cli>Completeness, revolving the outcome&rsquo;s certainty and omission-absent.\u003C/li>\r\n\u003Cli>Consistency, ensuring uniformity without discrepancies or contradictions.\u003C/li>\r\n\u003Cli>Conformity, verifying alignment with the anticipated format, standards, and structures.\u003C/li>\r\n\u003Cli>Accuracy, validating precision and correctness of the retrieved web intelligence.\u003C/li>\r\n\u003Cli>Integrity, reflecting unauthorized alterations capable of affecting the structured datasets.\u003C/li>\r\n\u003Cli>Timeliness, guaranteeing that the extracted info remains up-to-date and pertinent.\u003C/li>\r\n\u003C/ul>\r\n\u003Cp>Considering the following steps moves a performer closer to the ideal output.\u003C/p>\r\n\u003Ch3>\u003Ca name=\"anchor3\">\u003C/a>1. Choosing a framework and libraries\u003C/h3>\r\n\u003Cp>The preferred computing language may vary. Simple and prompt Ruby suits for small-scaled tasks, and C++ may be more optimal than CGI scripts. \u003Ca href=\"https://dexodata.com/en/blog/using-java-for-data-scraping-and-harvesting-on-the-web\" target=\"_blank\" rel=\"noopener\">Applying Java for online data harvesting\u003C/a> leads to fast info processing, and so on. As Python remains the most common scraping solution with a vast range of open-source libraries, we will concentrate on leveraging this language for gathering high-quality data. The selection of web parser stays at developer&rsquo;s discretion, as well as buying residential and mobile proxies, or datacenter ones.\u003C/p>\r\n\u003Cp>The Scrapy framework offers a flexible approach for CSS, HTML, PHP, and Node.js-oriented sites. Here is a basic Python script for retrieving lists of cities and their population from the target source without pagination and excluding a new project&rsquo;s creation:\u003C/p>\r\n\u003C/li>\r\n\u003C/ul>\r\n\u003Ctable style=\"border-collapse: collapse; width: 99.9794%;\" border=\"2\">\r\n\u003Ctbody>\r\n\u003Ctr>\r\n\u003Ctd style=\"width: 97.0536%;\">\r\n\u003Cp style=\"margin-top: 32px; font-weight: 400; padding-left: 40px;\">\u003Ccode>\u003Cspan style=\"color: #2e95d3;\">import&nbsp;\u003C/span>scrapy\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>class&nbsp;\u003Cspan style=\"color: #f22c3d;\">CityPopulationSpider\u003C/span>(scrapy.Spider):\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; name =&nbsp;\u003Cspan style=\"color: #00a67d;\">'city_population'\u003C/span>\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; start_urls = [\u003Cspan style=\"color: #00a67d;\">'https://site-to-scrape.com/cities'\u003C/span>]&nbsp;\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp;&nbsp;\u003Cspan style=\"color: #2e95d3;\">def\u003C/span>&nbsp;\u003Cspan style=\"color: #f22c3d;\">parse\u003C/span>(self, response):\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;\u003Cspan style=\"color: #7e8c8d;\"># Replace 'city_selector' and 'population_selector' with the actual HTML selectors, and site-to-scrape.com with actual address\u003C/span>\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; city_elements = response.css(\u003Cspan style=\"color: #00a67d;\">'div.city_selector'\u003C/span>)\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp;&nbsp;\u003Cspan style=\"color: #2e95d3;\">for&nbsp;\u003C/span>city_element&nbsp;\u003Cspan style=\"color: #2e95d3;\">in&nbsp;\u003C/span>city_elements:\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; city_name = city_element.css(\u003Cspan style=\"color: #00a67d;\">'span.name::text'\u003C/span>).get()\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; population = city_element.css(\u003Cspan style=\"color: #00a67d;\">'span.population::text'\u003C/span>).get()\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;\u003Cspan style=\"color: #2e95d3;\">yield&nbsp;\u003C/span>{\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;\u003Cspan style=\"color: #00a67d;\">&nbsp;'city_name'\u003C/span>: city_name,\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;\u003Cspan style=\"color: #00a67d;\">&nbsp;'population'\u003C/span>: population,\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; }\u003C/code>\u003C/p>\r\n\u003C/td>\r\n\u003C/tr>\r\n\u003C/tbody>\r\n\u003C/table>\r\n\u003Cp>The corrected script puts collected elements into the &ldquo;citiesandpopulation.json&rdquo; file after running:\u003C/p>\r\n\u003Ctable style=\"border-collapse: collapse; width: 99.9794%;\" border=\"2\">\r\n\u003Ctbody>\r\n\u003Ctr>\r\n\u003Ctd style=\"width: 97.0536%; padding-left: 40px;\">\u003Cspan style=\"font-family: monospace; font-weight: 400; background-color: #b4d7ff;\">scrapy crawl city_population -o citiesandpopulation.json\u003C/span>\u003C/td>\r\n\u003C/tr>\r\n\u003C/tbody>\r\n\u003C/table>\r\n\u003Cp>&nbsp;\u003C/p>\r\n\u003Ch3>\u003Ca name=\"anchor4\">\u003C/a>2. Crawling sites\u003C/h3>\r\n\u003Cp>Navigation between the same site&rsquo;s sections is called pagination, while crawling is the same process applied to multiple web pages. To optimize the work with numerous sources, a reliable ecosystem of geo targeted proxies is applied. Ethical intermediaries distribute the load on servers and assist in avoiding throttling, the excess of queries per time unit. Scrapy as a fast tool serves for crawling, and the basic script for the &ldquo;site-to-scrape.com&rdquo; example looks like:\u003C/p>\r\n\u003Ctable style=\"border-collapse: collapse; width: 99.9794%;\" border=\"2\">\r\n\u003Ctbody style=\"padding-left: 40px;\">\r\n\u003Ctr style=\"padding-left: 40px;\">\r\n\u003Ctd style=\"width: 97.0536%; padding-left: 40px;\">\r\n\u003Cp style=\"margin-top: 32px; font-weight: 400; padding-left: 40px;\">\u003Ccode>scrapy startproject job_crawler\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>\u003Cspan style=\"color: #f22c3d;\">cd&nbsp;\u003C/span>job_crawler\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>scrapy genspider example site-to-scrape.com\u003C/code>\u003C/p>\r\n\u003C/td>\r\n\u003C/tr>\r\n\u003C/tbody>\r\n\u003C/table>\r\n\u003Cp>&nbsp;\u003C/p>\r\n\u003Ch3>\u003Ca name=\"anchor5\">\u003C/a>3. Running dynamic proxies\u003C/h3>\r\n\u003Cp>Harvesting top-notch web insights at corporate scale requires one to buy residential and mobile proxies in sufficient amounts. \u003Ca href=\"https://dexodata.com/en/blog/static-and-dynamic-proxies-everything-you-need-to-know-about\" target=\"_blank\" rel=\"noopener\">Dynamic servers changing external addresses\u003C/a> within a previously set IP pool are now common. They ensure continuous data gathering. Libraries like \u003Ccode>scrapy-proxies\u003C/code> or \u003Ccode>requests \u003C/code>perform authorization for every IP, change address, and repeat the cycle. In the following example for the \u003Ccode>requests \u003C/code>library, \u003Ccode>HTTPProxyAuth \u003C/code>manages the access stage based on login-password entering:\u003C/p>\r\n\u003Ctable style=\"border-collapse: collapse; width: 99.9794%;\" border=\"2\">\r\n\u003Ctbody style=\"padding-left: 40px;\">\r\n\u003Ctr style=\"padding-left: 40px;\">\r\n\u003Ctd style=\"width: 97.0536%; padding-left: 40px;\">\r\n\u003Cp style=\"margin-top: 32px; font-weight: 400; padding-left: 40px;\">\u003Ccode>\u003Cspan style=\"color: #2e95d3;\">import&nbsp;\u003C/span>requests\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>\u003Cspan style=\"color: #2e95d3;\">from&nbsp;\u003C/span>requests.auth&nbsp;\u003Cspan style=\"color: #2e95d3;\">import&nbsp;\u003C/span>HTTPProxyAuth\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>\u003Cspan style=\"color: #2e95d3;\">from\u003C/span>\u003Cspan style=\"color: #2e95d3;\">&nbsp;\u003C/span>itertools&nbsp;\u003Cspan style=\"color: #2e95d3;\">import&nbsp;\u003C/span>cycle\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Cspan style=\"color: #7e8c8d;\">\u003Ccode># Insert proper IP, port, and authentication details below\u003C/code>\u003C/span>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Cspan style=\"color: #7e8c8d;\">\u003Ccode>proxies_list = [\u003C/code>\u003C/span>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; {\u003Cspan style=\"color: #00a67d;\">'http'\u003C/span>:&nbsp;\u003Cspan style=\"color: #00a67d;\">'http://username1:password1@proxy1:port1'\u003C/span>, '\u003Cspan style=\"color: #00a67d;\">https'\u003C/span>:\u003Cspan style=\"color: #00a67d;\">&nbsp;'http://username1:password1@proxy1:port1'\u003C/span>},\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; {\u003Cspan style=\"color: #00a67d;\">'http'\u003C/span>:&nbsp;\u003Cspan style=\"color: #00a67d;\">'http://username2:password2@proxy2:port2'\u003C/span>,&nbsp;\u003Cspan style=\"color: #00a67d;\">'https'\u003C/span>:&nbsp;\u003Cspan style=\"color: #00a67d;\">'http://username2:password2@proxy2:port2'\u003C/span>},\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp;&nbsp;\u003Cspan style=\"color: #7e8c8d;\"># Add more servers as needed\u003C/span>\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>]\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode># Set up authorization\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>proxy_auth_list = [HTTPProxyAuth(proxy[\u003Cspan style=\"color: #00a67d;\">'http'\u003C/span>].split(\u003Cspan style=\"color: #00a67d;\">'@'\u003C/span>)[\u003Cspan style=\"color: #f22c3d;\">0\u003C/span>].split(\u003Cspan style=\"color: #00a67d;\">'://'\u003C/span>)[\u003Cspan style=\"color: #f22c3d;\">1\u003C/span>].split(\u003Cspan style=\"color: #00a67d;\">':'\u003C/span>)[\u003Cspan style=\"color: #f22c3d;\">0\u003C/span>],\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; proxy[\u003Cspan style=\"color: #00a67d;\">'http'\u003C/span>].split(\u003Cspan style=\"color: #00a67d;\">'@'\u003C/span>)[\u003Cspan style=\"color: #f22c3d;\">0\u003C/span>].split(\u003Cspan style=\"color: #00a67d;\">'://'\u003C/span>)[\u003Cspan style=\"color: #f22c3d;\">1\u003C/span>].split(\u003Cspan style=\"color: #00a67d;\">':'\u003C/span>)[\u003Cspan style=\"color: #f22c3d;\">1\u003C/span>])\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;\u003Cspan style=\"color: #2e95d3;\">for&nbsp;\u003C/span>proxy&nbsp;\u003Cspan style=\"color: #2e95d3;\">in\u003C/span>&nbsp;proxies_list]\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Cspan style=\"color: #7e8c8d;\">\u003Ccode># Here is a cycle&rsquo;s example for dynamic geo targeted proxies\u003C/code>\u003C/span>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>proxy_cycle = cycle(proxies_list)\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>auth_cycle = cycle(proxy_auth_list)\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>def make_request(url):\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Cspan style=\"color: #7e8c8d;\">\u003Ccode>&nbsp; &nbsp; # Continue creating authenticated proxy pairs\u003C/code>\u003C/span>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; current_proxy = next(proxy_cycle)\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; current_auth = next(auth_cycle)\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; try:\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; response = requests.get(url, proxies=current_proxy, auth=current_auth, timeout=10\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Cspan style=\"color: #7e8c8d;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; # Process the response as needed\u003C/code>\u003C/span>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; print(f\"Proxy: {current_proxy}, Status Code: {response.status_code}\")\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; except Exception as e:\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; print(f\"Error: {e}\")\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Cspan style=\"color: #7e8c8d;\">\u003Ccode>&nbsp; &nbsp; &nbsp; &nbsp; # Handle errors if needed\u003C/code>\u003C/span>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Cspan style=\"color: #7e8c8d;\">\u003Ccode># Example usage\u003C/code>\u003C/span>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>url_to_scrape =&nbsp;\u003Cspan style=\"color: #00a67d;\">'https://site-to-scrape.com'\u003C/span>\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>\u003Cspan style=\"color: #2e95d3;\">for&nbsp; &nbsp;in&nbsp;\u003Cspan style=\"color: #f22c3d;\">range\u003C/span>\u003C/span>(\u003Cspan style=\"color: #f22c3d;\">5\u003C/span>):&nbsp;&nbsp;\u003Cspan style=\"color: #7e8c8d;\"># Make 5 requests (example value) using different proxies\u003C/span>\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp; make_request(url_to_scrape)\u003C/code>\u003C/p>\r\n\u003C/td>\r\n\u003C/tr>\r\n\u003C/tbody>\r\n\u003C/table>\r\n\u003Cp>&nbsp;\u003C/p>\r\n\u003Ch3>\u003Ca name=\"anchor6\">\u003C/a>4. Cleaning and preprocessing raw data\u003C/h3>\r\n\u003Cp>The raw material needs cleaning and preprocessing to raise the quality. Some values miss or duplicate during the initial scraping phase, others differ significantly from the main targets (outliers). Cleaning includes numerous tweaks, such as:\u003C/p>\r\n\u003Col>\r\n\u003Cli>Converting categorical variables to numerals accordingly\u003C/li>\r\n\u003Cli>Normalizing metrics to average weights\u003C/li>\r\n\u003Cli>Turning the current features to new ones for their clarification, especially in AI-based scraping techniques\u003C/li>\r\n\u003Cli>\u003Ca href=\"https://dexodata.com/en/blog/data-enrichment-with-geo-targeted-proxies-general-overview\" target=\"_blank\" rel=\"noopener\">Enriching data through geo targeted proxies\u003C/a> with additional elements.\u003C/li>\r\n\u003C/ol>\r\n\u003Cp>The \u003Ccode>pandas \u003C/code>library in Python is a reliable instrument. It can delete duplicates and standardize date formats in a few steps, as shown here:\u003C/p>\r\n\u003Ctable style=\"border-collapse: collapse; width: 99.9794%;\" border=\"2\">\r\n\u003Ctbody>\r\n\u003Ctr>\r\n\u003Ctd style=\"width: 97.0536%;\">\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>\u003Cspan style=\"color: #2e95d3;\">import&nbsp;\u003C/span>pandas&nbsp;\u003Cspan style=\"color: #2e95d3;\">as&nbsp;\u003C/span>pd\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>df = pd.read_csv(\u003Cspan style=\"color: #00a67d;\">'raw_data.csv'\u003C/span>)\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>df = df.drop_duplicates()\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>df[\u003Cspan style=\"color: #00a67d;\">'date'\u003C/span>] = pd.to_datetime(df[\u003Cspan style=\"color: #00a67d;\">'date'\u003C/span>],&nbsp;\u003Cspan style=\"color: #f22c3d;\">format\u003C/span>=\u003Cspan style=\"color: #00a67d;\">'%Y-%m-%d'\u003C/span>)\u003C/code>\u003C/p>\r\n\u003C/td>\r\n\u003C/tr>\r\n\u003C/tbody>\r\n\u003C/table>\r\n\u003Cp>&nbsp;\u003C/p>\r\n\u003Ch3>\u003Ca name=\"anchor7\">\u003C/a>5. Applying machine learning\u003C/h3>\r\n\u003Cp>AI-driven models capable of processing natural language are a common support tool for setting objectives and writing scripts, e.g. Copilot or \u003Ca href=\"https://dexodata.com/en/blog/chatgpt-and-data-collection-implications-for-web-data-harvesting-software\" target=\"_blank\" rel=\"noopener\">ChatGPT assisting in data extraction\u003C/a>. Machine learning models pass training to obtain relevant information from unstructured assets, such as text or images. In Python, the \u003Ccode>spaCy \u003C/code>library is responsible for deploying ML-oriented logics. Training a multi-layered AI-enhanced neural network for gathering online info at scale is a complicated task. The primary entities retrieval from publicly open sites via&nbsp;\u003Cspan style=\"font-family: monospace; background-color: #b4d7ff;\">spaCy \u003C/span>however takes such forms, in case of targeting on names and locations from news feeds:\u003C/p>\r\n\u003Ctable style=\"border-collapse: collapse; width: 99.9794%;\" border=\"2\">\r\n\u003Ctbody>\r\n\u003Ctr>\r\n\u003Ctd style=\"width: 97.0536%;\">\r\n\u003Cp style=\"margin-top: 32px; font-weight: 400; padding-left: 40px;\">\u003Ccode>\u003Cspan style=\"color: #2e95d3;\">import&nbsp;\u003C/span>spacy\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>nlp = spacy.load(\u003Cspan style=\"color: #00a67d;\">'en_core_web_sm'\u003C/span>)\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>text =&nbsp;\u003Cspan style=\"color: #00a67d;\">\"The ethical Dexodata ecosystem offered to buy residential IP pools located in San-Francisco, Beijing, Paris, and 100+ countries&rsquo; locations.\"\u003C/span>\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>doc = nlp(text)\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>\u003Cspan style=\"color: #2e95d3;\">for\u003C/span>&nbsp;ent&nbsp;\u003Cspan style=\"color: #2e95d3;\">in&nbsp;\u003C/span>doc.ents:\u003C/code>\u003C/p>\r\n\u003Cp style=\"font-weight: 400; padding-left: 40px;\">\u003Ccode>&nbsp; &nbsp;&nbsp;\u003Cspan style=\"color: #f22c3d;\">print\u003C/span>(\u003Cspan style=\"color: #00a67d;\">f'{ent.text}: {ent.label_}'\u003C/span>)\u003C/code>\u003C/p>\r\n\u003C/td>\r\n\u003C/tr>\r\n\u003C/tbody>\r\n\u003C/table>\r\n\u003Cp>The next stage requires recurrent data cleaning and preprocessing within common AI models&rsquo; training.\u003C/p>\r\n\u003Ch3>\u003Ca name=\"anchor8\">\u003C/a>High-quality web data collection and Dexodata\u003C/h3>\r\n\u003Cp>Gathering high-quality web data at scale faces ethical considerations in addition to maintaining sustainable link with target sites. Buying residential and mobile proxies from the Dexodata ecosystem solves these issues. \u003Ca href=\"https://dexodata.com/en/blog/why-dexodata-implements-aml-and-kyc-policies\" target=\"_blank\" rel=\"noopener\">Dexodata acts in strict compliance with KYC and AML policies\u003C/a> when it comes to acquring and supporting IP addresses. Applying our geo targeted proxies in conjunction with adherence to sites&rsquo; terms of service and robots.txt rules, avoiding overburdening, and respecting copyright, will lead your business to supreme internet data for further business development. Reach out to our support for getting a free proxy trial.\u003C/p>\u003C/body>\u003C/html>\n",[16,19,21,24,27,30,33],{"lang":17,"slug":18},"ru","izvlecenie-dannyx-s-veb-stranic-v-korporativnyx-masstabax-5-bazovyx-sagov-s-python-i-dexodata",{"lang":20,"slug":6},"en",{"lang":22,"slug":23},"ua","ua-izvlecenie-dannyx-s-veb-stranic-v-korporativnyx-masstabax-5-bazovyx-sagov-s-python-i-dexodata",{"lang":25,"slug":26},"cn","cn-gathering-high-quality-web-data-at-scale-5-tweaks-to-apply",{"lang":28,"slug":29},"es","es-gathering-high-quality-web-data-at-scale-5-tweaks-to-apply",{"lang":31,"slug":32},"fr","fr-gathering-high-quality-web-data-at-scale-5-tweaks-to-apply",{"lang":34,"slug":35},"ar","ar-gathering-high-quality-web-data-at-scale-5-tweaks-to-apply",[37,38],"Data collection","Guide",[],1784902993207]