
In Progress
Posted
Paid on delivery
I need a robust web-scraping solution that can reliably pull 100,000 text-based records from a set of e-commerce sites. The data I’m after is purely textual—product titles, prices, descriptions, categories and any other publicly visible details that help build a clean catalogue. No images are required. Here’s how I envision the job: • You create an automated scraper (Python, Scrapy, BeautifulSoup, Selenium or a comparable stack) that navigates the target stores, handles pagination, variations and search filters, and respects reasonable crawl rates while bypassing blocks or CAPTCHAs when they appear. • The script should be reusable: site list, category paths and output format must be easy for me to tweak later. • Final deliverables: the fully commented source code, a requirements file or environment export, and a CSV or JSON file containing at least 100,000 unique product rows with the agreed-upon fields. • Acceptance criteria: random spot-checks of 200 rows show <2 % missing values per field; duplicate rate under 1 %; script completes a full run on a fresh machine with a single command. Let me know which libraries you prefer, any rate-limiting strategy you plan to adopt, and how long you’ll need to hit the 100k target.
Project ID: 40651487
53 proposals
Remote project
Active 14 hours ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs

I can build a reusable Python scraper using Scrapy/Playwright, with pagination, filters, variations, deduplication, and proper rate limiting. I’ll structure it so the target sites and output fields are easy to update, and deliver 100,000+ clean unique records, source code, requirements file, and CSV/JSON. I can also handle common anti-bot challenges within reasonable limits. I’d estimate 3–5 days depending on the target sites.
$20 CAD in 5 days
3.7
3.7
53 freelancers are bidding on average $28 CAD for this job

Hello, As a seasoned developer with more than 15 years of experience and an in-depth understanding of web scraping and automation, I have dealt with numerous challenges similar to your project. My proficiency in using powerful languages and tools like Python, Scrapy, BeautifulSoup, Selenium, and more, positions me well to handle the required scraping project for your e-commerce site. With every project I undertake, my priority is always on creating reusable, scalable solutions that can be easily tweaked later if needed - exactly what your project envisions. Having extensively worked with WordPress, Shopify, and Laravel among other technologies that power e-commerce sites like the ones you want me to scrape - I tend to stay updated with the latest security protocols and am experienced in dealing with roadblocks such as CAPTCHAs or blocks during scraping. Ultimately, my aim would be to deliver a fully commented source code that accommodates any future changes you may have and a final product that meets your exact needs. I know you also need speed without compromising on quality. This is something I'm capable of delivering. I will implement an effective rate-limiting strategy to ensure smooth process and synchronization across varied websites. With me on the job you can breathe easy knowing that <2% missing values per field and under 1% duplicate rate are guaranteed. Plus, I'm confident I can complete this task in a timely fashion whi Thanks!
$30 CAD in 1 day
7.9
7.9

By hiring BN-Droids Digital Services, you're entrusting your crucial data extraction project to a seasoned professional team. Our expertise lies in developing automated web scraping solutions using cutting-edge libraries like Python, Scrapy, BeautifulSoup, and Selenium. As a result, we deliver advanced, reusable scripts that bypass blocks and CAPTCHAs while maintaining reasonable crawl rates - ensuring efficient completion of the scraping task without compromising on data accuracy.
$30 CAD in 1 day
7.0
7.0

Have over 18 years of experience in data mining/ Web scrapping/ Scraping Bots/ Chrome/Opera Extensions I have done it all. Tell us your source and we will put it in excel for you, Or we can even give you filtered results as per your requirement, In the format you want. You can also ask for data into a particular format - Excel, Json, Mysql, Databases, XMLs, you name them. Further Can help you with integrating it with ur databases, Can create json outputs. We are not only good with scraping but also with the tools that u may need after that. We can help you build you softwares round the data we have 99% Data Accuracy. We have Duplicate finder. etc., We can help with Statistics on the data We can help with creating Api's front the data We can create Softwares to manage that data We can build Sites round the data
$120 CAD in 2 days
6.9
6.9

Dear client, I can do this job of E-commerce Text Data Scraping accurately as per your requirements and available to start immediately. Thanks!
$30 CAD in 1 day
6.7
6.7

Hi there, I can create a robust web scraper using Python, leveraging Scrapy and BeautifulSoup to reliably gather 100,000 textual records from your specified e-commerce sites. I'll ensure the scraper handles pagination, variations, and search filters effectively while respecting crawl rates. Additionally, I will include safeguards to bypass blocks or CAPTCHAs as necessary. To meet your requirements, I'll provide thoroughly commented source code, along with a requirements file for easy modifications. The output will be structured to meet your data validation criteria, ensuring both efficiency and reliability. Your satisfaction is my priority and I guarantee that I will deliver you a high-quality result. Regards, Ali
$30 CAD in 1 day
6.3
6.3

Hi, I will build a Python scraper using Scrapy with Selenium fallback to pull 100,000 unique product rows (titles, prices, descriptions, categories) into CSV or JSON, with reusable site list and configurable output. I can start today. For rate limits, I will use rotating proxies, randomized delays, and retry logic to stay under block thresholds while keeping duplicates below 1%. Questions: 1) Which specific e-commerce sites are the targets? 2) CSV or JSON preferred for the final dataset? Looking forward to discussing further. Regards, Shayan.
$11 CAD in 1 day
5.3
5.3

Hi, If the budget is flexible, I’d use Scrapy with async requests and rotating user agents for speed, plus a deduplication pipeline keyed on normalized title + category to hit your <1% duplicate target. Rate limiting would be adaptive per domain (2–5 req/sec) to avoid blocks while maintaining throughput. But at current pricing, delivering reusable, reliable code with spot-checkable data quality isn’t feasible without cutting corners that violate your specs. Is there room to adjust the budget to reflect actual effort, or should we scope down to a smaller pilot first? Share one target site URL so I can estimate realistic extraction effort and confirm whether 100k is achievable within a viable budget.
$30 CAD in 1 day
5.4
5.4

I can build a reusable Python/Scrapy scraping pipeline to collect 100,000+ unique e-commerce product records with titles, prices, descriptions, categories and other public text fields. Approach: Scrapy for high-volume crawling, BeautifulSoup/lxml for parsing, Selenium only where JavaScript rendering is necessary, plus throttling, retries, proxy/session handling and duplicate validation. For sites with CAPTCHAs or access controls, I’d use compliant alternatives where possible rather than relying on brittle bypasses.
$30 CAD in 1 day
5.4
5.4

Hi! My name is Marjan and I'm here to offer you my services as a skilled applicant with over a decade of experience working on Freelancer.com. l believe I am the best fit candidate for this project due to my extensive experience; I would like to have a discussion to get to know that we both are on the same page. Once the scope will be locked, I will start working on it right away.
$200 CAD in 7 days
5.5
5.5

Hello there. I hope you are donig well. I have successfully developed robust web scraping solutions for various e-commerce platforms, efficiently extracting product data like titles, prices, and descriptions. My proficiency in Python and libraries such as Scrapy and BeautifulSoup ensures that I can deliver high-quality results tailored to your specific needs. I understand that you require a reliable scraper to collect 100,000 text-based records while managing pagination and bypassing blocks. I will implement an automated solution that adheres to best practices, ensuring minimal downtime and high data quality through efficient error handling and rate-limiting strategies. I will provide a fully documented source code, a requirements file, and a structured CSV or JSON output with at least 100,000 unique product entries. My approach focuses on quality and performance, ensuring that the final product meets your acceptance criteria with a low missing and duplicate rate. Please feel free to reach out to me. I look forward to working with you. Best regards, Billy Bryan
$18 CAD in 5 days
4.8
4.8

Hello, "Scrapy Pipeline With Deduplication + Rate Limits" - you need 100k clean product rows without the scraper becoming fragile. I’d use Scrapy for the main crawl, with BeautifulSoup only where page parsing needs it. I’d separate site/category settings from the crawler so you can change targets later without rewriting the scraper. I’d also normalize fields and deduplicate by stable product identifiers or canonical URLs before export. For blocks, I’d use conservative concurrency, delays, retries and backoff rather than aggressive requests. The final run will be tested from a clean environment and checked against your 2% missing-value and 1% duplicate requirements. Which e-commerce sites and approximate number of categories are included in the 100k target? Looking forward to working with you. Truong
$20 CAD in 1 day
4.9
4.9

★•══•★ Hi client ★•══•★ I’m confident I can build a solid scraper that pulls clean, text-only product data from your target sites. I’d use Python with Scrapy for efficiency and flexibility, adding Selenium only if JavaScript-heavy pages or CAPTCHAs pop up. I’ll handle pagination, filters, and rate limiting carefully to avoid blocks—think polite crawling with randomized delays. The script will be easy to tweak for new sites or categories, and I’ll deliver well-commented code plus a CSV or JSON with 100k+ unique products. Spot checks will ensure data quality stays tight with minimal missing info or duplicates. Ready to dive into this and get you a reliable tool that just works. What’s your preferred way to share the site list and category paths? Best regards, Rico
$20 CAD in 7 days
5.0
5.0

Hello ? Your project is interesting because the main focus is clean and reusable data extraction, not just a one-time scrape. I work with Python and data processing, and I can build a structured scraping solution that collects product titles, prices, descriptions, categories, and exports them to CSV or JSON. I’d also keep the configuration simple so site lists and output fields can be updated later without changing the whole script. For large datasets like 100,000 records, I would use a modular approach with pagination handling, duplicate checks, and rate limiting to improve stability. Price: 30 CAD fixed Delivery: 5 days Could you share the target websites or a few example URLs? That will help me estimate the scraping method and confirm the timeline more accurately. Best regards, Albert C.
$30 CAD in 5 days
4.9
4.9

AVAILABLE TO START IMMEDIATELY..,,. I will deliver a robust, reusable Python (Scrapy/Selenium) scraper for 100,000 e-commerce text records (titles, prices, descriptions), along with commented source code and a CSV/JSON file. 10+ years Advanced Excel experience, Certified VBA Programmer, MBA.
$19 CAD in 1 day
4.9
4.9

I have done data extraction jobs like this before, pulling 100k plus records without getting blocked or losing fields halfway through. I would build it in Python with Scrapy and rotate through Selenium only where the site needs JS rendering, output straight to CSV or a database. Can start today. First working batch in 2 to 3 days. The 10 to 30 CAD and timeline are based on what's in the post. Once we go over which sites and fields you need, I will firm both up. Want me to send a quick scope doc?
$30 CAD in 5 days
4.0
4.0

The 100k-row target combined with a $10–30 CAD budget and 4-day window is the real design constraint here — at that price the only way to hit six-figure volume is a highly parallelized, resumable crawler rather than a single long-running script, since a single interruption on a fresh machine would blow the "one command, full run" acceptance criterion. I'd build it in Scrapy for the concurrent request handling and pagination/category-path logic, use Playwright only on pages that need JS rendering or filter interaction, and checkpoint progress to SQLite so a crashed run resumes instead of restarting, then export to CSV/JSON at the end with a dedup pass keyed on product URL to hit your <1% duplicate target. Rate limiting would be adaptive per-domain (AutoThrottle plus backoff on 429/CAPTCHA responses) rather than a fixed delay, so it stays under block thresholds without wasting time on sites that tolerate faster crawling. Which sites are you targeting — are they all similarly structured, or does the scraper need to handle meaningfully different layouts per store, since that changes how much of the "reusable" config layer needs building versus per-site custom parsing? Send the site list and field list and I'll start.
$10 CAD in 4 days
3.7
3.7

Hello, I reviewed your project regarding **E-commerce Text Data Scraping**, and I understand that you need a robust web-scraping solution that can reliably collect 100,000 clean text-based product records from multiple e-commerce sites. I have over three years of experience with **Python, JavaScript, Web Scraping, Data Extraction, and Selenium** and have worked on similar projects involving reusable crawlers, pagination handling, and structured CSV/JSON output. I can provide a reliable, maintainable, and professional solution tailored to your requirements. For this project, I will: • Review the requirements and current setup • Implement **an automated scraper with configurable site paths, filters, and output formats** • Test the solution thoroughly • Provide clear progress updates • Deliver clean, documented, and maintainable work • Support you with any related issues after delivery I am available to begin immediately and can complete the work within **10 days**. Please share any existing files, references, or additional requirements so I can confirm the best implementation approach. Which e-commerce sites are in scope, and do any of them already have predictable pagination or anti-bot hurdles to plan around? Best regards, Miguel
$100 CAD in 10 days
3.2
3.2

Hi! there - Abror here "100,000 PRODUCT RECORDS" — you need a reusable scraper that stays clean and reliable across pagination, variations, and changing store layouts. I’d use Scrapy for the main crawl, with structured item validation and duplicate checks before writing CSV/JSON. I’d keep site URLs, category paths and output fields in config so you can change targets without touching the scraper logic. For blocks, I’d use sensible request rates, retries and session handling rather than aggressive crawling. CAPTCHA or access controls should be handled according to the target site's rules, not blindly bypassed. Which e-commerce sites are the targets, and do they have roughly similar page structures?
$20 CAD in 1 day
3.2
3.2

At 100,000 records the risk isn't writing the scraper, it's keeping it running without getting blocked halfway through. I'd build it in Scrapy with rotating delays and retry logic per site, structure the output schema up front (title, price, description, category) so it's consistent across sources, and validate a sample against the live pages before the full run. You'd get the scraper code plus the clean dataset. How many distinct e-commerce sites are in scope, and do they share a similar page structure or are they all different?
$18 CAD in 4 days
2.9
2.9

Here’s a concise proposal you can send: Hi, I can build a reusable Python scraping system capable of collecting 100,000+ product records while keeping the code modular and easy to extend. I’d use Scrapy + Playwright/Selenium where JavaScript rendering is actually required, with configurable site/category inputs, pagination and variation handling, structured extraction, normalization, and deduplication. My approach includes: Product title, price, description, category and agreed public fields Configurable spiders/selectors per website Pagination, filters and product variations Request throttling, retries and respectful crawl rates Data validation and duplicate detection CSV/JSON export Logging and error reporting Requirements file + one-command setup Clean, commented and reusable source code For anti-bot protections, I can implement legitimate retry/throttling and browser-rendering strategies, but I would not bypass CAPTCHAs or access controls. I’ll validate the final dataset against your <2% missing-field and <1% duplicate acceptance criteria and provide the 100,000+ unique records. If you share the target websites and required fields, I can confirm the exact scraping approach and timeline before starting.
$15 CAD in 1 day
2.3
2.3

Montreal, Canada
Payment method verified
Member since May 30, 2026
$250-750 CAD
$30-250 CAD
$10-30 CAD
₹12500-37500 INR
$250-750 USD
₹1500-12500 INR
$1500-3000 AUD
$30-250 AUD
$30-250 USD
₹12500-37500 INR
₹600-1500 INR
₹1500-12500 INR
$8-15 USD / hour
₹750-1250 INR / hour
₹1500-12500 INR
₹2000-4000 INR
₹12500-37500 INR
$250-750 USD
₹12500-37500 INR
₹1500-12500 INR
₹600-601 INR
$15-25 USD / hour
₹12500-37500 INR