Skip to content
insyncrg-bitPublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

Startup onboarding OCR service

Extracts structured information for a memo from:

  1. LinkedIn profile – Open Graph meta tags (og:title, og:image, og:description) for name and profile picture.
  2. Pitch deck – Text from PDF or PPTX, then heuristic section extraction (problem, solution, market, team, traction, ask, etc.).

LinkedIn and deck parsing are separate; you can call each on its own or combine via the runner. Output is a single passable JSON suitable for a future API.

Setup

cd ocr-service
python -m venv .venv
source .venv/bin/activate   # or .venv\Scripts\activate on Windows
pip install -r requirements.txt

Usage

LinkedIn only:

python run.py --linkedin "https://www.linkedin.com/in/username"

Deck only (PDF or PPTX):

python run.py --deck path/to/pitch_deck.pdf
python run.py --deck path/to/pitch_deck.pptx

Both (full memo):

python run.py --linkedin "https://www.linkedin.com/in/username" --deck path/to/deck.pdf

Write JSON to file:

python run.py -l "https://linkedin.com/in/foo" -d deck.pdf -o memo.json

Output shape (passable JSON)

  • linkedin: url, name, profile_pic, description, raw_og, error
  • deck: full DeckExtract (company_name, tagline, problem, solution, sections, raw_slides, source_type, error)
  • memo: merged fields for the memo (founder_name, founder_photo_url, company_name, tagline, problem, solution, market, traction, team, ask, sections, raw_slides, etc.)

Programmatic use

from linkedin_parser import parse_linkedin_profile
from deck_parser import parse_pdf_deck, parse_pptx_deck, parse_deck
from run import build_memo

# LinkedIn only
profile = parse_linkedin_profile("https://linkedin.com/in/username")
print(profile.name, profile.profile_pic)

# Deck only (PDF and PPTX are separate functions)
deck_pdf = parse_pdf_deck("pitch.pdf")
deck_pptx = parse_pptx_deck("pitch.pptx")
# Or auto-dispatch by extension:
deck = parse_deck("pitch.pdf")

# Full memo JSON
memo = build_memo(
    linkedin_url="https://linkedin.com/in/username",
    deck_path="pitch.pdf",
)

Notes

  • LinkedIn: Profile pages may require login in a browser; OG tags are often still present in the initial HTML for crawlers. Use a browser-like User-Agent; for production you may need a headless browser or proxy if results are empty.
  • PDF: Uses pdfplumber (layout-aware text extraction).
  • PPTX: Uses python-pptx; .ppt (binary) is not supported.
  • Structured extraction: Rule-based section detection (keywords like "Problem", "Solution", "Traction"); no LLM. You can later add an SLM/API step on the extracted text for richer parsing.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages