Make historic House Statements of Disbursements (~1980–present) machine-readable by extracting, cleaning, and normalizing PDFs into a consistent, analyzable dataset.
Problem
The House’s Statements of Disbursements (spending reports) are primarily available as scanned PDF documents, making data analysis and historical trend identification extremely difficult.
Solution
- Digitize and structure historic Statements of Disbursements data, going back to approximately 1980, into machine-readable spreadsheets.
- Data Source: Utilize existing scrapers developed by the Sunlight Foundation and ProPublica, supplemented by Optical Character Recognition (OCR) and data cleaning techniques. Source historical documents from the Boston Public Library (and other potential archives).
- Data Structure: Create a consistent schema for the extracted data, allowing for easy querying and analysis across different years and reporting periods.