
Closed
Posted
Paid on delivery
I have a batch of scanned-PDF files—together they hold somewhere between 11 and 50 individual tables—but the challenge is that each table follows its own layout. I need every one of those tables accurately lifted out of the scans, cleaned, and delivered in a single Excel workbook. OCR will definitely be required; a mix of tools such as Adobe, ABBYY, Tabula, Camelot, or even custom Python scripts is fine as long as the final spreadsheet is precise and ready for analysis. What matters to me: • Accuracy: numbers and text must match the originals, column headings preserved, no merged-cell surprises. • Consistency: even though the table structures differ, each finished sheet in Excel should be formatted in a tidy, readable way. • Organisation: name each worksheet logically or consolidate tables with clear separators—whatever keeps the data easy to navigate. • Turnaround: let me know how fast you can complete the job once you’ve seen the files; speed is appreciated but quality comes first. I’ll provide the PDFs as soon as we agree. Before hand-off, I’ll spot-check random rows against the scans; if everything lines up, we’re done. Looking forward to seeing how you tackle this.
Project ID: 40655062
44 proposals
Remote project
Active 3 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
44 freelancers are bidding on average ₹19,181 INR for this job

I can complete this project by combining OCR extraction (Tesseract, ABBYY-style layout OCR, Camelot/pdfplumber) with manual visual verification against each scan, since tables with differing layouts need different extraction methods rather than one blanket approach — this avoids merged-cell errors and misread characters (0/O, 1/l/I, decimals). Every extracted table gets cross-checked against the source image before entering Excel, not just validated by OCR confidence scores. In the final workbook, each table gets its own worksheet named for its source file/page or subject matter, formatted consistently with bold frozen headers, auto-fit columns, borders, and professional fonts, plus a summary index sheet listing every table and its source for easy navigation. All numbers are stored as true numeric values, not text, so the sheet is ready for formulas and analysis immediately. Turnaround depends on table complexity and scan quality: straightforward printed tables can be done same-session, a mixed batch of 20–35 tables likely same day, and 40–50 tables or heavily degraded scans will need a firm estimate once I see them, since I won't guess pricing or timing blind. Send the PDFs and I'll confirm an exact turnaround before starting, then deliver the workbook with a self-run spot-check already done so your review is confirmation rather than first-pass debugging.
₹20,000 INR in 2 days
5.5
5.5

Your biggest risk is OCR drift on multi-layout tables—if column boundaries shift by even two pixels between scans, you'll end up with misaligned data that breaks pivot tables downstream. I've built Python pipelines that handle exactly this: adaptive bounding-box detection plus post-OCR validation rules that flag anomalies before they corrupt your workbook. Quick questions - are any of these tables nested or do they span multiple pages? And do you need formulas preserved if the originals contain calculated cells? Here is the architectural approach: - PYTHON + TABULA/CAMELOT: Parse structured tables first, fall back to ABBYY FineReader for low-quality scans, then run regex validators on numeric columns to catch OCR errors. - DATA PROCESSING: Normalize headers across sheets using fuzzy-match logic so "Revenue Q1" and "Q1 Revenue" map to one standard column name. - EXCEL DELIVERY: Export to .xlsx with conditional formatting on suspect cells, plus a summary sheet showing confidence scores per table so you know where to spot-check. I've extracted 200+ tables from legacy insurance PDFs where every form had different grid layouts. Let's do a 10-minute call after you share two sample files so I can scope turnaround and flag any edge cases now.
₹22,500 INR in 7 days
5.4
5.4

★★★ TOP 1% IN FREE LANCER WORLD ★★★ ★★★ 20+ Year Experience in IBD being CMD★★★ ★★★ 200+ Country Satisfied Clientele ★★★ ★Linkedin★ ★Data Entry★ ★Business Plans★★★ ★★★ Operational Strategic planner Customer Support 24*7★★★ ★★★Excel/Word Operation★★★ ★★★Chat Support★★★ ★★★Calling Support★★★ ★★★Business Plans / Marketing Strategy ★★★ * Digital Marketing★★★ ★★★Social Media Marketing ★★★ ★★★Internet Marketing ★★★ ★★★Any type of Data Projects★★★ ★★★★★★ Regards, ★★★CMD★★★ ★★★PVSYS GROUP (INDIA)★★★ ★★★IF YOU THINK THEN I CAN★★★
₹37,500 INR in 99 days
5.2
5.2

Hi, I can extract, clean, and structure all 11–50 tables from your scanned PDF batch into a single, analysis-ready Excel workbook with 100% data fidelity. Technical Approach & Extraction Workflow Multi-Engine OCR & Parsing Pipeline: For varied table layouts, standard single-pass tools often cause misaligned columns and unwanted merged cells. I combine ABBYY FineReader Pro, Camelot / Tabula (both lattice & stream modes), and custom Python (pdfplumber, pytesseract, OpenCV) scripts to clean scan artifacts and preserve distinct table boundaries. Data Cleansing & Normalization: Eliminate scan artifacts, ensure proper numeric/date formatting, preserve multi-tier column headers, and eliminate accidental merged cells. Workbook Organization: Deliver a single, tidy Excel workbook with each table placed on a logically named sheet (or consolidated with distinct metadata headers and visual separators for easy navigation). Two-Step Quality Assurance: Automated checksum verifications (column sums matched against scan totals) paired with a manual line-by-line spot-check against the original PDFs before final delivery. Turnaround & Estimated Timeline 11–25 Tables: 24–36 hours from file delivery. 25–50 Tables: 48–72 hours from file delivery. Please feel free to share the PDFs or a 1–2 page sample; I would be glad to provide a quick sample extraction for your verification.
₹25,000 INR in 7 days
4.4
4.4

I am Mahad Sheikh, an experienced technology professional focused on creating efficient solutions to complex problems - data processing being one of my specialties. I have a solid understanding and skills in Python and using tools such as Adobe, ABBYY, Tabula, Camelot to extract and organize data from various file formats including scanned PDFs. This project aligns perfectly with my expertise and the challenges you highlighted are ones I have tackled successfully in the past. Accuracy and consistency of output are paramount in this project and I assure you that my clients often commend me for these qualities. I can harness the power of OCR and use customized Python scripts to ensure alignment of numbers, text and column headers. This will ensure no unpleasant surprises like merged cells, preserving the original content while optimizing it for readability. Every worksheet in the Excel workbook will be logical named or appropriately divided for easy navigation and analysis. Time management is not just a buzzword for me; it's key to delivering quality work promptly. As usual, quality is my priority but after analyzing the files you send, I'll let you know an estimated turnaround time - always striving for a balance between speed and precision. Your assurance of spot-checking random rows against scans tells me you value a thorough job; with me on your project, that is exactly what you'll get. So, let's get started on this intricate task together
₹12,500 INR in 5 days
4.2
4.2

Since these are scanned PDFs with 11 to 50 tables that each follow a different layout, treating this as one generic OCR job would misfire - I'd start with a quick layout audit to group tables by structure, then use Tabula/Camelot for the clean grids and fall back to ABBYY or a custom OCR+regex pass for anything with merged cells or multi-row headers, which is usually where automated extraction actually breaks. Every sheet comes out with clean headers and consistent formatting, not a raw dump, and I'll run my own row-by-row check against the scans before handoff so your spot-check comes back clean. Roughly how many of the tables have irregular layouts like merged cells versus straightforward grids - that shapes how I split automated vs manually-verified extraction.
₹13,500 INR in 5 days
3.9
3.9

Hi. I can extract and clean all tables from your scanned PDFs. I can finish in a couple of days. Majority time will be analyzing the table formats so we can identify the table and extract the information. I will use custom Python scripts combining OCR tools like Tesseract or pdfplumber with pandas to handle varying table layouts, clean merged cells, and organize data into clean worksheets. I am available to start immediately. Reach out so we can discuss the files and get started. cheers Nehal
₹15,000 INR in 2 days
3.5
3.5

Hi, I can extract all tables from your scanned PDF files and deliver them in one clean Excel workbook with accurate text, numbers, headings, and readable formatting. My approach will be to first review a sample PDF and identify table layouts, scan quality, OCR difficulty, and output structure. Then I’ll use the best mix of OCR tools such as ABBYY/Adobe plus manual verification, and Python tools if useful, to recreate each table cleanly in Excel. I’m comfortable with scanned PDF table extraction, OCR cleanup, Excel formatting, data validation, numeric checking, multi-layout tables, and final quality review. Deliverables: * Single Excel workbook * All 11–50 tables extracted * Original headings preserved * Text and numbers verified * Separate logical worksheets * Clean readable formatting * No unwanted merged cells * Random row cross-checks * Final accuracy review Timeline: 3–5 days after reviewing the PDFs, depending on scan quality and table complexity. I’ll focus on accuracy first, with careful side-by-side checking so the final workbook is ready for analysis and passes your spot checks. Best regards Ankit
₹12,500 INR in 4 days
3.7
3.7

Greetings Sir/Madam. Hope you are doing great. I am an experienced professional in web scraping / OCR/ Data accumulation and ready to work on your project. Rest assured that you will be delivered quality service in a very short amount of time at a very reasonable rate. Feel free to contact me to discuss further details. Thanks in advance
₹12,500 INR in 7 days
4.5
4.5

Hi, I will extract every table from your scanned PDFs using OCR (ABBYY or Python with Tesseract) and deliver one tidy Excel workbook with headings preserved and no merged-cell issues. I can start today. For each PDF I will use a separate worksheet with a logical name, then spot-check rows against the scans before hand-off. Questions: 1) Roughly how many PDF files in total? 2) Any preferred sheet naming, by source file or by table topic? Looking forward to discussing further. Regards, Shayan.
₹21,250 INR in 3 days
2.5
2.5

Hi, I read your project carefully and I can deliver exactly what you need. I can start immediately and show you the first preview in a few hours. Let’s discuss the details.
₹12,500 INR in 4 days
2.4
2.4

Your requirement is less about simple PDF conversion and more about building a reliable extraction workflow that can handle inconsistent table structures without losing data integrity. For scanned PDFs with varying layouts, I would approach this in stages: OCR preprocessing, table boundary detection, structured extraction, normalization, and validation before generating the final Excel workbook. For the OCR layer, I can combine tools such as ABBYY/Tesseract with Python-based processing depending on scan quality. For extraction, I typically use a mix of Camelot/Tabula and custom parsing scripts to handle irregular columns, broken rows, merged cells, and inconsistent headers. The important part is not only extracting the tables, but validating that totals, numeric formatting, and column alignment match the original scans. The final workbook will be organized for practical use: clean worksheet naming, consistent formatting, preserved headers, and readable layouts even when source tables differ significantly. I also include verification checks during processing to reduce OCR-related mismatches before delivery. Once I review a sample PDF batch, I can confirm the exact turnaround time and extraction strategy. Based on the described volume, I expect delivery within a few days while maintaining careful quality control for the spot checks you mentioned.
₹31,320.96 INR in 4 days
2.3
2.3

Hello, I can extract all tables from your scanned PDFs using OCR combined with Python-based processing and table extraction tools. I’ll preserve the original headings, wording, numbers, and column structure while cleaning the output into a readable Excel workbook. Each table will be placed on a logically named worksheet, with formatting kept consistent and no unnecessary merged cells. I’ll also manually/automatically cross-check extracted values against the scans before delivery. **Turnaround:** 24–48 hours after receiving the PDFs, depending on the number and complexity of tables. I can provide a small sample first so you can verify the extraction quality before I process the complete batch.
₹12,500 INR in 1 day
1.8
1.8

Hello, i read your requirement. I have experience in excel and done many projects i give you best work on your time and budget. Thanks, waiting for your response...
₹13,000 INR in 4 days
1.9
1.9

You're dealing with 50+ tables that all look different, so a one-size OCR tool will fail on half of them. I'd start by running each PDF through Adobe's OCR to get a baseline, then use Python with OpenCV to detect table borders and Camelot to extract the cleanest candidates. For the messy ones, I'll manually adjust the region of interest and re-run the extraction until the numbers match the scan. In my experience, most people stop at the first OCR pass and call it done, which is why spreadsheets end up with misaligned columns or missing headers. I'll add a validation step that compares row counts and checksums against the original PDF, so you get a report showing which tables needed extra attention. The final Excel workbook will have each table on its own sheet, named after the PDF filename and page number, with consistent formatting and no merged cells.
₹25,000 INR in 5 days
1.9
1.9

Your files are scans, so Camelot or Tabula alone will not hold up on them: those need a real text layer with ruling lines. The path that works here is OCR first (ABBYY or Tesseract with layout detection), then a parser profile per layout, since you said every table has its own structure. One profile per layout, not one generic pass that quietly mangles the odd table. How I would run it: 1. You send 2 or 3 representative pages. I return those tables as sheets before you commit to anything, so you judge accuracy on your own scans instead of on a promise. 2. Full batch: each table into its own named worksheet in one workbook, headings preserved, no merged cells, numbers stored as numbers so they are ready for analysis. 3. Verification pass before hand-off: row and column counts per table, plus every numeric column re-checked against the scan. Cells the OCR was unsure about get flagged rather than silently guessed, so your spot-check finds nothing. Where accuracy usually dies on scans is faint print, stamps running across cells and rotated or skewed pages. If your batch has those, say so and I will size the job honestly rather than discover it mid-way. About me: one completed project on this account, rated 5 out of 5, delivered on time and on budget. Outside the platform, 15 merged pull requests into third party open source projects, most of them into a Go security tool with 184 stars, each one reviewed and accepted by its maintainer. 14000 INR, 4 working days counted from the moment I have the files. If the batch turns out closer to 50 tables than 11 and the layouts differ wildly, you hear that from me before I start, not after. Petro Pankov, BotCraft Group
₹14,000 INR in 4 days
1.5
1.5

This project immediately caught my attention because it is exactly the type of work I do best. I understand you need accurate extraction of complex tables from scanned PDFs, ensuring each layout is cleaned and delivered as a single Excel workbook. Your emphasis on accuracy, consistency, and organization aligns perfectly with my skills. While I am new to freelancer, I have tons of experience and have done other projects off site. I have worked with OCR tools like Adobe and custom Python scripts to ensure precision and readability in data presentation. If this sounds like what you're looking for I'd love to hear more about your project. Regards, Warrick Van Eeden
₹16,900 INR in 7 days
1.2
1.2

As an expert in Adobe Acrobat, I have extensive experience in dealing with complex PDF files similar to the ones you currently have. My forte lies in extracting and cleaning up data from varying formats and presenting it error-free and organized, thus producing a final output that matches your criteria for accuracy, consistency, organization, and tight scheduling.
₹13,332 INR in 1 day
0.9
0.9

Hi, "each table follows its own layout" is the detail that rules out a one-size-fits-all tool - Tabula and Camelot work well on consistently-structured tables, but with 11-50 tables each structured differently, I'd expect to build a Python-based pipeline (pdfplumber/Camelot for structured extraction, Tesseract or a cloud OCR API for genuinely scanned/image-only pages) with per-table validation rather than assume one method handles everything cleanly. Process: OCR/extract each table individually, verify against the source scan (not just trust the extraction), clean merged-cell artifacts and preserve column headings exactly, then consolidate into one Excel workbook with clearly named worksheets or consistent separators - formatted tidily even though source structures vary. Given you'll spot-check rows against the scans before sign-off, I'll build in my own verification pass first so surprises are caught before delivery, not during your check. Turnaround estimate once I see the actual files - table complexity and scan quality vary enough that I'd rather give you an honest number after review than guess blind. Ready to receive the PDFs whenever you're ready. Arun
₹12,500 INR in 7 days
0.0
0.0

Hi — this is exactly the messy-scan extraction I do well. Since these are scans with a different layout per table, straight table-parsers (Camelot/Tabula) won't cut it alone — those need a digital text layer these files don't have. My pipeline: • OCR — run the scans through a strong engine (ABBYY / PaddleOCR) with table-structure recognition, so cells map to the right rows and columns even when layouts differ table to table. • RECONSTRUCT — rebuild each table in Python (pandas), preserving column headings and fixing any merged-cell or shifted-column artefacts by hand where the OCR is unsure. • VERIFY — I check every table against the source scan before hand-off, so it passes your random-row spot-check first time. • DELIVER — one Excel workbook, one tidy, well-formatted sheet per table, named logically for easy navigation. Send the PDFs and I'll confirm the exact turnaround once I see the layouts. ₹15,000, and I can move fast — quality first, as you said. — Aakaash
₹15,000 INR in 4 days
0.0
0.0

Koraput, India
Member since May 6, 2026
₹750-1250 INR / hour
$10-25 USD
₹750-1250 INR / hour
$8-15 USD / hour
₹400-750 INR / hour
₹12500-37500 INR
₹100-400 INR / hour
$15-25 USD / hour
$15-25 USD / hour
$15-25 USD / hour
$10-30 USD
₹750-1250 INR / hour
$30-250 USD
₹1500-12500 INR
$250-750 USD
$1500-3000 USD
£10-20 GBP
₹750-1250 INR / hour
₹600-1500 INR
₹100-400 INR / hour