Web Scraping Python: A Practical Guide With Proxies
ColdProxy Team7 min read

Web scraping with Python is now a standard way for developers and data teams to collect public web data at scale, especially for ecommerce price monitoring and retail intelligence. Automated data collection now underpins competitive analysis, ad verification, and geo-targeted testing across global markets. For developers new to the workflow, Python's requests and BeautifulSoup libraries handle static HTML parsing well on their own. Real-world price monitoring usually needs more: proxies to spread requests across many IPs and reduce per-IP rate-limit pressure. This guide covers building a working scraper for public retail data, adding residential proxies for geo-targeted checks, and picking a plan that fits your workload.
Key Takeaways
- Configure ColdProxy residential proxies in Python requests using a proxy dictionary with credentials and gateway details copied from your client area, not example values.
- Add geo-targeting tags, such as a country tag, directly to the proxy username, and copy the exact format from your client area because tag prefixes can differ between plans.
- Use sticky sessions, set with a numeric session tag and a time tag, for multi-page tests on sites you are authorized to test, and rely on default IP rotation for broad public price collection across retailers.
- Choose Residential IPv4 (GB Based) for traffic-budgeted projects or Residential IPv4 (Unmetered) Mbps speed tiers for ongoing ecommerce monitoring; both offer up to 0.6s response time across 195+ countries.
- Switch to Playwright headless browsers only when target pages require JavaScript execution to load price data, not as a default replacement for requests and BeautifulSoup.
Building a Python Price Scraper With Residential Proxies
You can build a working price scraper with just the requests library and BeautifulSoup to parse static HTML from public retail sites or practice environments. Install both packages with pip, then write a script that fetches a page, selects product elements with CSS selectors, and extracts titles and prices into CSV or JSON. A good practice target is books.toscrape.com, which lists book titles and prices on every page. That site exists for learning, so you can test selectors and pagination logic without risking blocks on live retail targets.
The ColdProxy residential proxy pool is ethically sourced and held to strict compliance standards.
When your target site serves different prices by location, or your collection volume puts pressure on a single IP, add proxies to your requests session to reduce per-IP rate-limit pressure. Configure the proxy dictionary with your credentials and gateway details exactly as shown below. The gateway and port here are examples; replace them with the values from your client area.
import requests
proxy = "http://USERNAME:PASSWORD@gw-2312.coldproxy.com:30000"proxies = {"http": proxy, "https": proxy}
resp = requests.get("https://api.vipv6proxy.com/api/checker/my-ip", proxies=proxies, timeout=30)print(resp.json()) # confirms the exit IPWith username and password authentication, you set geo-targeting for price monitoring in the username string. To request an IP from a specific country, add a country tag to the username. Tag prefixes can differ between plans, so copy the exact format from your client area; the username tag scheme documentation and the Python setup guide explain how the tags work. ColdProxy runs a 70M+ global Residential IPv4 pool across 195+ countries, with up to 0.6s response time.
Sticky sessions hold the same pool IP for a set window. That matters for multi-page flows on sites you are authorized to test, such as your own store's checkout, where session continuity counts for more than IP rotation. To hold an IP for a set window, such as ten minutes, add both a numeric session tag and a time tag to your username, as shown in the Python setup guide. The sticky vs rotating session differences guide walks through when to use broad rotation for collection versus a sticky session for multi-step testing. Without both tags, requests rotate IPs by default, spreading traffic across the network to reduce IP blocks during large collection runs.
None of this bypasses CAPTCHAs, gets around bot protection, or promises that every request will succeed against a well-defended retail target. You still need to respect robots.txt, build in reasonable rate limiting, and stay away from login-gated content or personal data. Check our ethical web scraping best practices and our web scraping proxies page before putting any collector into production.
Common failure points in Python scrapers: CSS selectors that break silently after a site redesign, missing rate limits that trigger blocks, and proxy authentication nobody checked before scaling up request volume. If you hit connection errors during testing, the connection reset troubleshooting guide helps you tell network issues, authentication failures, and target-side blocks apart before you start rewriting your script.
When to Use Headless Browsers for JavaScript Rendering
Requests and BeautifulSoup can't execute JavaScript. Pages that load price data dynamically through client-side rendering will hand your parser empty or incomplete HTML. Switch to a headless browser like Playwright only once you've confirmed, by inspecting the target page, that the content you need loads after JavaScript runs rather than in the initial server response. Our headless browser fundamentals guide covers how headless browsers work and the performance trade-offs behind that decision.
Playwright uses the same ColdProxy gateway, username and password, but takes them as separate proxy settings rather than one URL. Browser automation also adds overhead that rules it out as a default for high-volume price monitoring. The Playwright setup documentation covers configuration for headless browsers specifically. Save the tool for targets where static parsing genuinely fails, rather than reaching for it first. Many public retail product pages still serve enough HTML for requests-based extraction, so confirm the JavaScript dependence is real before adding that complexity to your pipeline.
Choosing the Right Proxy Plan for Ecommerce Monitoring
Choosing between ColdProxy's Residential IPv4 plans comes down to whether your monitoring workload is better budgeted by traffic volume or by throughput. Residential IPv4 (GB Based) sells monthly traffic packs and suits smaller jobs where you can estimate the total gigabytes a project needs. Residential IPv4 (Unmetered) sells by Mbps speed tier, with throughput capped at the chosen tier, and fits larger or ongoing monitoring work where predictable bandwidth matters more than counting gigabytes. Both plans draw on the same 195+ country pool with identical targeting; they differ in billing model, billing periods and activation time.
Current starting prices and plan structures are on the pricing page; no need to contact sales or request a quote. Residential IPv4 (GB Based) pricing shows monthly pack costs for teams budgeting by traffic. Residential IPv4 (Unmetered) pricing shows rates for each Mbps speed tier, with hourly to monthly billing. Both plans support HTTP, HTTPS, and SOCKS5 (UDP via SOCKS5), with rotating sessions or sticky sessions of 5 seconds to 24 hours.
Need a different speed or traffic allowance? Contact us through the channels on our website so we can assess whether a custom plan is possible.
Real-time retail price tracking needs fast, consistent response times. ColdProxy states up to 0.6s response time for both Residential IPv4 plans. Test against your own targets during the shortest paid billing window before committing to a longer term. There's no free trial, so treat your first purchase as a validation phase: measure success rates, latency, and geo-targeting accuracy against what your production setup actually needs.
Frequently Asked Questions
How do I configure ColdProxy residential proxies in a Python requests script?
Create a proxy dictionary with http and https keys that both point to one proxy URL built from your username, password, gateway host and port, as in the code example above. Pass the dictionary as the proxies argument of your requests call, and replace every example value with the details from your ColdProxy client area.
What username format enables country targeting and sticky sessions?
Add a country tag to your username for location-specific exits. For a sticky session, add both a numeric session tag and a time tag; a ten-minute time value holds the IP for ten minutes. Tag prefixes can differ between plans, so copy the exact format from your client area rather than guessing.
When should I switch from requests to Playwright for scraping?
Switch only when inspecting the target page confirms that price data loads via JavaScript after initial HTML delivery, making static parsing with BeautifulSoup insufficient. Many retail product pages serve adequate HTML for requests-based extraction, so validate JavaScript dependence before adding browser automation overhead.
Which ColdProxy plan fits small versus large-scale price tracking?
Residential IPv4 (GB Based) suits smaller projects budgeted by monthly traffic volume, while Residential IPv4 (Unmetered) fits ongoing monitoring where predictable Mbps throughput matters more than per-GB cost tracking. Both plans access the same 195+ country pool with identical targeting and up to 0.6s response time.
Is web scraping hard to learn for beginners?
Basic scraping with Python requests and BeautifulSoup is accessible for developers familiar with Python fundamentals, especially using practice sites like books.toscrape.com. The complexity increases when adding proxy integration, handling pagination, respecting rate limits, and debugging selector failures after site updates, but each skill builds incrementally through hands-on testing.


