Extracts structured information for a memo from:
- LinkedIn profile – Open Graph meta tags (
og:title,og:image,og:description) for name and profile picture. - Pitch deck – Text from PDF or PPTX, then heuristic section extraction (problem, solution, market, team, traction, ask, etc.).
LinkedIn and deck parsing are separate; you can call each on its own or combine via the runner. Output is a single passable JSON suitable for a future API.
cd ocr-service
python -m venv .venv
source .venv/bin/activate # or .venv\Scripts\activate on Windows
pip install -r requirements.txtLinkedIn only:
python run.py --linkedin "https://www.linkedin.com/in/username"Deck only (PDF or PPTX):
python run.py --deck path/to/pitch_deck.pdf
python run.py --deck path/to/pitch_deck.pptxBoth (full memo):
python run.py --linkedin "https://www.linkedin.com/in/username" --deck path/to/deck.pdfWrite JSON to file:
python run.py -l "https://linkedin.com/in/foo" -d deck.pdf -o memo.jsonlinkedin:url,name,profile_pic,description,raw_og,errordeck: fullDeckExtract(company_name, tagline, problem, solution, sections, raw_slides, source_type, error)memo: merged fields for the memo (founder_name, founder_photo_url, company_name, tagline, problem, solution, market, traction, team, ask, sections, raw_slides, etc.)
from linkedin_parser import parse_linkedin_profile
from deck_parser import parse_pdf_deck, parse_pptx_deck, parse_deck
from run import build_memo
# LinkedIn only
profile = parse_linkedin_profile("https://linkedin.com/in/username")
print(profile.name, profile.profile_pic)
# Deck only (PDF and PPTX are separate functions)
deck_pdf = parse_pdf_deck("pitch.pdf")
deck_pptx = parse_pptx_deck("pitch.pptx")
# Or auto-dispatch by extension:
deck = parse_deck("pitch.pdf")
# Full memo JSON
memo = build_memo(
linkedin_url="https://linkedin.com/in/username",
deck_path="pitch.pdf",
)- LinkedIn: Profile pages may require login in a browser; OG tags are often still present in the initial HTML for crawlers. Use a browser-like User-Agent; for production you may need a headless browser or proxy if results are empty.
- PDF: Uses
pdfplumber(layout-aware text extraction). - PPTX: Uses
python-pptx;.ppt(binary) is not supported. - Structured extraction: Rule-based section detection (keywords like "Problem", "Solution", "Traction"); no LLM. You can later add an SLM/API step on the extracted text for richer parsing.