Про автоматизацию набора книг через ИИ
Наверное, я бы не скоро ещё добрался до этого скрипта, но моя коллега меня к этому подтолкнула. Дело в том, что я ленивый, и книжки, которые перевожу, набираю не я, а моя коллега, Александра.
Я ей скидываю книги, которые мне присылает издательство на перевод, говорю что-то вроде: «Александра, я через пару дней буду пустой», что в переводе означает, что у меня заканчиваются подготовленные тексты на перевод, и она мне высылает свои наработки по набору.
И вот три дня назад пишу ей эту заветную фразу, а она отвечает: никак, уехала ненадолго в другой город, буду у компьютера через несколько дней. Ну что ж. Допиливаю остатки. И вот я реально пустой: либо набирать самому, либо ждать.
Набирать самому — совершенно неохота. Распознавание через Finereader — это редкостный геморрой: сначала чистишь текст от артефактов распознавания, потом форматируешь. А у меня ещё в работе книжка по математике — та самая, которую я иногда упоминаю в заметках, — с многоэтажными и длинными формулами вот такого вида:
Такое даже файнридер не берет.
Но незадолго до всех этих событий я экспериментировал с DeepSeek: скормил ему копипасту из этих формул. Он их понял и вывел в LaTeX. Я поделился наблюдениями с Александрой — и она таким макаром подготовила мне несколько глав. Тоже, впрочем, муторно: в книге почти 750 страниц, обалдеешь столько копипастить и склеивать.
ЭТАП 1: ПЕРВЫЙ СКРИПТ
И я решил этот опыт автоматизировать. Взял книгу по программированию на Go — текст там примерно такого характера:
За несколько итераций навайбкодил скрипт. Он умеет следующее:
- принимает путь к PDF-файлу плюс начальную и конечную страницы;
- вытаскивает текст не просто копипастой, а с метатегами: где курсив, где жирный, где заголовок, где код;
- чтобы не поперхнуться объемом и не жрать ресурсы, работает батчами по три странички;
- дальше соединяется с DeepSeek.Chat по API и через промпт просит преобразовать копипасту с метаинформацией в Markdown-файл;
- DeepSeek принимает под козырёк и выдаёт результат.
Два часа отладки — и три минуты работы скрипта. На выходе примерно 200 страниц набранного и отформатированного текста в Маркдауне:
Вот, собственно, и код на Python:
# %% [I. Imports]
import fitz # PyMuPDF
import json
import os
import re
import base64
from openai import OpenAI
from PIL import Image
import io
print("✅ Cell 1: Imports loaded successfully.")
# ==============================================================================
# %% [CONFIGURATION CONSTANTS]
# ==============================================================================
PDF_PATH = r"H:\...\GoBook.pdf"
START_PAGE = 170
END_PAGE = 374
PAGES_PER_CHUNK = 3
DEEPSEEK_API_KEY = "sk-..."
PROGRESS_FILE = "progress.json"
EXTRACT_IMAGES = True
MIN_IMAGE_SIZE = 10000
IMAGES_OUTPUT_FILE = "document_images_tables.md" # Explicit filename
print("✅ Cell 2: Configuration constants set.")
# ==============================================================================
# %% [API CLIENT INITIALIZATION]
# ==============================================================================
client = OpenAI(
api_key=DEEPSEEK_API_KEY,
base_url="https://api.deepseek.com"
)
print("✅ Cell 3: DeepSeek API client initialized.")
# ==============================================================================
# %% [SYSTEM PROMPTS]
# ==============================================================================
SYSTEM_PROMPT_TEXT = """
You are an expert assistant specialized in converting raw PDF text into clean, perfectly formatted Markdown.
CRITICAL RULES:
1. MERGE BROKEN LINES: Remove arbitrary line breaks within paragraphs. Merge lines to form continuous, flowing prose. ONLY use newlines for actual paragraph breaks, headings, lists, or code blocks.
2. HEADERS/FOOTERS: Aggressively DELETE running headers and footers. If you see a standalone page number (e.g., "164", "166") or a chapter title (e.g., "Chapter 7 Errors in Go") at the very beginning or end of a page chunk, REMOVE IT. It is not part of the main text.
3. CODE BLOCKS & CALLOUTS: PDF extraction often interleaves marginal notes (callouts) into the middle of code blocks.
- Reconstruct the code block as a SINGLE, valid Markdown code block (```go ... ```).
- Move any interrupting descriptive text (e.g., "**Function returning only an error**") to a bulleted list or blockquote IMMEDIATELY AFTER the code block.
4. HEADINGS & LISTINGS: Recognize patterns like "Listing 7.1" or "Chapter 7" and format them as Markdown headings (e.g., `### Listing 7.1: main.go`). Preserve `**bold**` formatting.
5. BULLETS: Convert weird bullet artifacts (like "¡", "", "•") into standard Markdown dashes (`- `).
6. NO SUMMARIZATION: Preserve 100% of the original information, logic, and technical depth.
7. NO META-TEXT: Output ONLY the cleaned Markdown. Do NOT output page markers like "[CONTEXT_MARKER: PAGE X]".
"""
SYSTEM_PROMPT_IMAGE = """
You are an expert OCR assistant specialized in extracting text from images and figures.
YOUR TASK:
1. Analyze the image and determine if it contains any text.
2. If text is present, extract it LINE BY LINE in the exact order it appears (top to bottom).
3. Format the output as a Markdown table with TWO columns:
- Column 1: "Original Line" - the exact text from the image
- Column 2: "Translation" - leave this column EMPTY (just a placeholder for future translation)
OUTPUT FORMAT EXAMPLE:
| Original Line | Translation |
|---------------|-------------|
| Compressor | |
| f: X\\to\\{0,1\\}^* | |
STRICT RULES:
- If NO text is detected, output exactly: "[NO_TEXT_DETECTED]"
- Do NOT describe the image. ONLY extract text.
- Preserve mathematical notation.
"""
print("✅ Cell 4: System prompts configured.")
# ==============================================================================
# %% [HELPER FUNCTIONS]
# ==============================================================================
def load_progress(pdf_path: str) -> int:
"""Loads progress ONLY if the PDF path matches. Prevents cross-file conflicts."""
if os.path.exists(PROGRESS_FILE):
try:
with open(PROGRESS_FILE, "r", encoding="utf-8") as f:
data = json.load(f)
if data.get("pdf_path") == pdf_path:
return data.get("last_processed_page", START_PAGE - 1)
except (json.JSONDecodeError, IOError):
pass
return START_PAGE - 1
def save_progress(pdf_path: str, page_num: int) -> None:
"""Saves progress with the PDF path to ensure file-specific tracking."""
with open(PROGRESS_FILE, "w", encoding="utf-8") as f:
json.dump({"pdf_path": pdf_path, "last_processed_page": page_num}, f)
def clean_pdf_artifacts(text: str) -> str:
"""Pre-processes text to fix math artifacts and weird bullet points."""
# Fix math
text = text.replace(" ∗", "^*").replace("∗", "^*")
text = text.replace("→", "\\to ").replace("←", "\\gets ")
# Fix weird bullets (inverted exclamation, dots, squares often found in PDFs)
text = re.sub(r'^\s*[¡•▪▫]\s*', '- ', text, flags=re.MULTILINE)
# Clean excessive spaces
text = re.sub(r' {2,}', ' ', text)
return text
def extract_formatted_text_from_pages(pdf_path: str, start: int, end: int) -> tuple:
"""Extracts text, merges lines to prevent mid-sentence breaks, and extracts images."""
doc = fitz.open(pdf_path)
text_chunks = []
images_data = []
for i in range(start - 1, end):
page_num = i + 1
page = doc[i]
text_chunks.append(f"\n[CONTEXT_MARKER: PAGE {page_num}]\n")
page_dict = page.get_text("dict")
page_text_parts = []
for block in page_dict.get("blocks", []):
if block.get("type") == 0: # Text block
block_lines = []
for line in block.get("lines", []):
line_parts = []
for span in line.get("spans", []):
text = span.get("text", "")
if not text.strip():
continue
flags = span.get("flags", 0)
is_italic = bool(flags & 2)
is_bold = bool(flags & 16)
text = clean_pdf_artifacts(text)
if is_bold and is_italic:
line_parts.append(f"***{text}***")
elif is_bold:
line_parts.append(f"**{text}**")
elif is_italic:
line_parts.append(f"*{text}*")
else:
line_parts.append(text)
# Join spans in a line with a space
block_lines.append(" ".join(line_parts))
# Join lines within a block with a SPACE, not a newline.
# This prevents mid-sentence line breaks. The LLM will handle paragraph splits.
page_text_parts.append(" ".join(block_lines))
text_chunks.append("\n\n".join(page_text_parts)) # Double newline separates distinct blocks
# Image extraction
if EXTRACT_IMAGES:
print(f" Scanning page {page_num} for images...")
image_list = page.get_images(full=True)
valid_images = []
for img_index, img_info in enumerate(image_list):
try:
xref = img_info[0]
base_image = doc.extract_image(xref)
img = Image.open(io.BytesIO(base_image["image"]))
if img.width * img.height >= MIN_IMAGE_SIZE:
valid_images.append((img, img_index))
except Exception:
continue
if valid_images:
print(f" 📎 Found {len(valid_images)} valid image(s) on page {page_num}")
for img, img_idx in valid_images:
print(f" 🖼️ Processing image {img_idx + 1}...")
try:
buffer = io.BytesIO()
img.save(buffer, format="PNG")
img_base64 = f"data:image/png;base64,{base64.b64encode(buffer.getvalue()).decode()}"
response = client.chat.completions.create(
model="deepseek-chat",
messages=[
{"role": "system", "content": SYSTEM_PROMPT_IMAGE},
{"role": "user", "content": [{"type": "image_url", "image_url": {"url": img_base64}}, {"type": "text", "text": "Extract text to table."}]}
],
max_tokens=1000,
temperature=0.1
)
extracted_text = response.choices[0].message.content.strip()
if extracted_text != "[NO_TEXT_DETECTED]":
images_data.append({"page": page_num, "image_index": img_idx, "extracted_text": extracted_text})
print(" ✅ Text extracted.")
else:
print(" ⚪ No text detected.")
except Exception as e:
print(f" ❌ Vision API error: {e}")
else:
print(f" ✓ No valid images found on page {page_num}")
doc.close()
return "".join(text_chunks).strip(), images_data
def validate_length(original_text: str, markdown_text: str) -> tuple[bool, str]:
"""Validates the output length."""
len_orig = len(original_text)
len_md = len(markdown_text)
if len_orig == 0: return True, "Empty original, skipping."
ratio = len_md / len_orig
if ratio < 0.8: return False, f"DATA LOSS! Ratio: {ratio:.2f}"
if ratio > 2.5: return False, f"HALLUCINATION! Ratio: {ratio:.2f}"
return True, f"Length OK (ratio: {ratio:.2f})"
print("✅ Cell 5: Helper functions defined.")
# ==============================================================================
# %% [MAIN PROCESSING LOOP]
# ==============================================================================
def process_pdf() -> None:
print("=" * 60)
print("🚀 STARTING PDF TO MARKDOWN PIPELINE")
print("=" * 60)
last_processed = load_progress(PDF_PATH)
current_start = last_processed + 1
if current_start > END_PAGE:
print("✅ All requested pages for THIS FILE have been processed.")
return
base_name = os.path.splitext(PDF_PATH)[0]
md_filename = f"{base_name}_clean.md"
print(f"📁 Text output: {md_filename}")
print(f"📁 Images output: {IMAGES_OUTPUT_FILE}\n")
all_images_data = []
while current_start <= END_PAGE:
current_end = min(current_start + PAGES_PER_CHUNK - 1, END_PAGE)
print("-" * 60)
print(f"🔄 Processing chunk: Pages {current_start} to {current_end}")
print("-" * 60)
print(" [1/4] Extracting text and images...")
raw_text, images_data = extract_formatted_text_from_pages(PDF_PATH, current_start, current_end)
all_images_data.extend(images_data)
if not raw_text.strip():
print(" ⚠️ Chunk empty. Skipping.")
current_start = current_end + 1
save_progress(PDF_PATH, current_end)
continue
print(" ✅ Extraction complete.")
print(" [2/4] Sending text to DeepSeek API...")
try:
response = client.chat.completions.create(
model="deepseek-chat",
messages=[
{"role": "system", "content": SYSTEM_PROMPT_TEXT},
{"role": "user", "content": raw_text}
],
temperature=0.1
)
markdown_result = response.choices[0].message.content.strip()
print(" ✅ Response received.")
except Exception as e:
print(f" ❌ API Error: {e}")
break
print(" [3/4] Validating output...")
is_valid, msg = validate_length(raw_text, markdown_result)
if not is_valid:
print(f" 🛑 HALTED: {msg}")
break
print(f" ✅ {msg}")
print(" [4/4] Saving results...")
with open(md_filename, "a", encoding="utf-8") as f:
f.write(markdown_result)
f.write("\n\n")
save_progress(PDF_PATH, current_end)
print(f" ✅ Saved and checkpointed: pages [{current_start}-{current_end}]")
current_start = current_end + 1
# Force save images file if any data exists
if all_images_data:
print("\n" + "=" * 60)
print(f"💾 Saving {len(all_images_data)} image table(s) to: {IMAGES_OUTPUT_FILE}")
print("=" * 60)
try:
with open(IMAGES_OUTPUT_FILE, "w", encoding="utf-8") as f:
f.write("# Extracted Text from Images/Figures\n\n")
f.write("Column 2 is intentionally left empty for translation.\n\n---\n\n")
for img_data in all_images_data:
f.write(f"## Page {img_data['page']}, Image {img_data['image_index'] + 1}\n\n")
f.write(img_data['extracted_text'])
f.write("\n\n---\n\n")
print("✅ Image tables saved successfully!")
except Exception as e:
print(f"❌ Failed to save images file: {e}")
else:
print("\n⚠️ No text-containing images were found in this run.")
print("=" * 60)
print("🎉 PIPELINE COMPLETED SUCCESSFULLY")
print("=" * 60)
if __name__ == "__main__":
# Clear old progress if you want to force re-run on the same file:
# if os.path.exists(PROGRESS_FILE): os.remove(PROGRESS_FILE)
process_pdf()
# %%ЭТАП 2: ПРОБЛЕМА МАТЕМАТИЧЕСКИХ ФОРМУЛ
Дальше я решил таким же макаром подготовить математическую книгу. Указал путь, попробовал на одной главе — и обломался. DeepSeek почему-то отформатировал переменные в формулах просто как жирные или курсивные буковки: просто поставил вокруг них одну или две звёздочки в разметке Маркдауна, без всякой LaTeX разметки. Но блин, он же LaTeX-разметку формул мне уже делал. Что пошло не так?
После некоторого дебаггинга выяснилось: модель может делать что-то одно. Получает текст с метатегами форматирования — и буквы в формулах форматирует. Получает просто текст без разметки — видит формулы, ставит LaTeX, но понятия не имеет, где у меня жирный, где курсив: в простом тексте эти маркеры теряются.
ЭТАП 3: РЕШЕНИЕ — ДВУХПОТОЧНЫЙ ПРОМПТ (DUAL-INPUT)
Выкурив пару рюмок кофе, я нашел решение: отдавать в модель две копипасты — простой текст и с метатегами.
- [RAW TEXT]: Простой текст без разметки. Модель использует его строго для идентификации математических формул, переменных и символов (чтобы понять, что "a 1 , a 2" — это
$a_1, a_2$).
- [FORMATTED TEXT]: Текст с правильной структурой абзацев, маркерами жирного/курсива и исправленными разрывами строк. Модель использует его как базу для структуры.
Дальше дело техники и часа вайбкодинга с тестированием. Потестил, получил главу с формулами и форматированием — и зарядил в скрипт.
Вот как это выглядит в набранном виде — это скриншот текста в Маркдауне + LaTeX того куска, что я скриншотил выше:
# %% [I. Imports]import
import fitz # PyMuPDF
import json
import os
import re
import base64
from openai import OpenAI
from PIL import Image
import io
print("✅ Cell 1: Imports loaded successfully.")
# ==============================================================================
# %% [CONFIGURATION CONSTANTS]
# ==============================================================================
PDF_PATH = r"H:\...\Information Theory.pdf"
START_PAGE = 225
END_PAGE = 713
PAGES_PER_CHUNK = 3
DEEPSEEK_API_KEY = "sk-..."
PROGRESS_FILE = "progress_math.json"
EXTRACT_IMAGES = True
MIN_IMAGE_SIZE = 10000
print("✅ Cell 2: Configuration constants set.")
# ==============================================================================
# %% [API CLIENT INITIALIZATION]
# ==============================================================================
client = OpenAI(
api_key=DEEPSEEK_API_KEY,
base_url="https://api.deepseek.com"
)
print("✅ Cell 3: DeepSeek API client initialized.")
# ==============================================================================
# %% [SYSTEM PROMPTS]
# ==============================================================================
SYSTEM_PROMPT_TEXT = """
You are an expert assistant specialized in converting PDF text into perfect Markdown with LaTeX math.
YOU WILL RECEIVE TWO VERSIONS OF THE SAME TEXT:
1. [RAW TEXT]: Unformatted, preserves original character sequence. Use this STRICTLY to identify mathematical formulas, variables, and symbols (e.g., recognizing that "a 1 , a 2" means "$a_1, a_2quot;).
2. [FORMATTED TEXT]: Contains correct paragraph structure, bold/italic markers, and fixed line breaks. Use this as the BASE for your output.
YOUR TASK:
Merge the best of both. Output a single, clean Markdown document that:
- Uses the paragraph structure, lists, and bold/italic formatting from [FORMATTED TEXT].
- Corrects ALL mathematical notation into proper LaTeX (e.g., `$a_1
Про автоматизацию набора книг через ИИ — Teletype
, `$S^n
Про автоматизацию набора книг через ИИ — Teletype
, `$X = (S_1, \\dots, S_n)
Про автоматизацию набора книг через ИИ — Teletype
) based on the context from [RAW TEXT].
- Aggressively removes running headers, footers, and isolated bold page numbers (e.g., "**196**", "**Part II**").
- Preserves 100% of the original information. NO SUMMARIZATION.
- Outputs ONLY the final Markdown. Do NOT output the raw text or any meta-commentary.
"""
SYSTEM_PROMPT_IMAGE = """
You are an expert OCR assistant. Extract text LINE BY LINE from the image.
Format as a 2-column Markdown table: "Original Line" and "Translation" (leave empty).
If NO text is detected, output exactly: "[NO_TEXT_DETECTED]". Do NOT describe the image.
"""
print("✅ Cell 4: System prompts configured.")
# ==============================================================================
# %% [HELPER FUNCTIONS]
# ==============================================================================
def load_progress() -> int:
if os.path.exists(PROGRESS_FILE):
try:
with open(PROGRESS_FILE, "r", encoding="utf-8") as f:
data = json.load(f)
if data.get("pdf_path") == PDF_PATH:
return data.get("last_processed_page", START_PAGE - 1)
except (json.JSONDecodeError, IOError):
pass
return START_PAGE - 1
def save_progress(page_num: int) -> None:
with open(PROGRESS_FILE, "w", encoding="utf-8") as f:
json.dump({"pdf_path": PDF_PATH, "last_processed_page": page_num}, f)
def extract_raw_text_for_context(pdf_path: str, start: int, end: int) -> str:
"""Extracts plain text without any markdown formatting, preserving original spacing for math context."""
doc = fitz.open(pdf_path)
text_chunks = []
for i in range(start - 1, end):
page = doc[i]
# "text" mode gives raw string, we just clean extreme artifacts
raw = page.get_text("text")
raw = raw.replace("fi", "fi").replace("fl", "fl")
text_chunks.append(raw)
doc.close()
return "\n".join(text_chunks).strip()
def merge_lines_smartly(block_lines: list) -> str:
if not block_lines:
return ""
merged = [block_lines[0].strip()]
for i in range(1, len(block_lines)):
prev_line = block_lines[i-1].strip()
curr_line = block_lines[i].strip()
if not curr_line:
continue
if re.search(r'[.!?]\s*#39;, prev_line):
merged.append("\n\n" + curr_line)
else:
merged.append(" " + curr_line)
return "".join(merged)
def remove_programmatic_headers_footers(text: str) -> str:
text = re.sub(r'\n\s*\*\*\d+\*\*\s*\n', '\n', text)
text = re.sub(r'\n\s*\*\*Part II\*\*\s*\n\s*\*\*Lossless Data Compression\*\*\s*\n', '\n', text, flags=re.IGNORECASE)
text = re.sub(r'\n{3,}', '\n\n', text)
return text.strip()
def extract_formatted_text_from_pages(pdf_path: str, start: int, end: int) -> str:
"""Extracts text with bold/italic markers and smart line merging."""
doc = fitz.open(pdf_path)
page_text_parts = []
for i in range(start - 1, end):
page = doc[i]
page_dict = page.get_text("dict")
block_texts = []
for block in page_dict.get("blocks", []):
if block.get("type") == 0:
block_lines = []
for line in block.get("lines", []):
line_parts = []
for span in line.get("spans", []):
text = span.get("text", "")
if not text.strip():
continue
# Fix ligatures early
text = text.replace("fi", "fi").replace("fl", "fl")
flags = span.get("flags", 0)
is_italic = bool(flags & 2)
is_bold = bool(flags & 16)
if is_bold and is_italic:
line_parts.append(f"***{text}***")
elif is_bold:
line_parts.append(f"**{text}**")
elif is_italic:
line_parts.append(f"*{text}*")
else:
line_parts.append(text)
block_lines.append(" ".join(line_parts))
merged_block = merge_lines_smartly(block_lines)
if merged_block:
block_texts.append(merged_block)
page_text_parts.append("\n\n".join(block_texts))
doc.close()
raw_formatted = "\n\n".join(page_text_parts)
return remove_programmatic_headers_footers(raw_formatted)
def extract_formatted_text_and_images(pdf_path: str, start: int, end: int) -> tuple:
"""Wrapper that gets formatted text and handles images."""
formatted_text = extract_formatted_text_from_pages(pdf_path, start, end)
images_data = []
if EXTRACT_IMAGES:
doc = fitz.open(pdf_path)
for i in range(start - 1, end):
page_num = i + 1
page = doc[i]
print(f" Scanning page {page_num} for images...")
image_list = page.get_images(full=True)
valid_images = []
for img_index, img_info in enumerate(image_list):
try:
xref = img_info[0]
base_image = doc.extract_image(xref)
img = Image.open(io.BytesIO(base_image["image"]))
if img.width * img.height >= MIN_IMAGE_SIZE:
valid_images.append((img, img_index))
except Exception:
continue
if valid_images:
print(f" 📎 Found {len(valid_images)} valid image(s) on page {page_num}")
for img, img_idx in valid_images:
print(f" 🖼️ Processing image {img_idx + 1}...")
try:
buffer = io.BytesIO()
img.save(buffer, format="PNG")
img_base64 = f"data:image/png;base64,{base64.b64encode(buffer.getvalue()).decode()}"
response = client.chat.completions.create(
model="deepseek-chat",
messages=[
{"role": "system", "content": SYSTEM_PROMPT_IMAGE},
{"role": "user", "content": [{"type": "image_url", "image_url": {"url": img_base64}}, {"type": "text", "text": "Extract text to table."}]}
],
max_tokens=1000,
temperature=0.1
)
extracted_text = response.choices[0].message.content.strip()
if extracted_text != "[NO_TEXT_DETECTED]":
images_data.append({"page": page_num, "image_index": img_idx, "extracted_text": extracted_text})
print(" ✅ Text extracted.")
except Exception as e:
print(f" ❌ Vision API error: {e}")
else:
print(f" ✓ No valid images found on page {page_num}")
doc.close()
return formatted_text, images_data
def validate_length(original_text: str, markdown_text: str) -> tuple[bool, str]:
len_orig = len(original_text)
len_md = len(markdown_text)
if len_orig == 0: return True, "Empty original, skipping."
ratio = len_md / len_orig
if ratio < 0.8: return False, f"DATA LOSS! Ratio: {ratio:.2f}"
if ratio > 2.5: return False, f"HALLUCINATION! Ratio: {ratio:.2f}"
return True, f"Length OK (ratio: {ratio:.2f})"
print("✅ Cell 5: Helper functions defined.")
# ==============================================================================
# %% [MAIN PROCESSING LOOP]
# ==============================================================================
def process_pdf() -> None:
print("=" * 60)
print("🚀 STARTING PDF TO MARKDOWN PIPELINE (DUAL-INPUT MODE)")
print("=" * 60)
last_processed = load_progress()
current_start = last_processed + 1
if current_start > END_PAGE:
print("✅ All requested pages for THIS FILE have been processed.")
return
base_name = os.path.splitext(PDF_PATH)[0]
md_filename = f"{base_name}_clean.md"
images_md_filename = f"{base_name}_images_tables.md"
print(f"📁 Text output: {md_filename}")
print(f"📁 Images output: {images_md_filename}\n")
all_images_data = []
while current_start <= END_PAGE:
current_end = min(current_start + PAGES_PER_CHUNK - 1, END_PAGE)
print("-" * 60)
print(f"🔄 Processing chunk: Pages {current_start} to {current_end}")
print("-" * 60)
print(" [1/4] Extracting DUAL text versions and images...")
raw_text = extract_raw_text_for_context(PDF_PATH, current_start, current_end)
formatted_text, images_data = extract_formatted_text_and_images(PDF_PATH, current_start, current_end)
all_images_data.extend(images_data)
if not formatted_text.strip():
print(" ⚠️ Chunk empty. Skipping.")
current_start = current_end + 1
save_progress(current_end)
continue
print(" ✅ Extraction complete.")
print(" [2/4] Sending DUAL-INPUT to DeepSeek API...")
try:
# Combine both texts into a single, structured prompt
combined_user_prompt = f"""
[RAW TEXT - Use for math context only]:
{raw_text}
---
[FORMATTED TEXT - Use as base structure]:
{formatted_text}
"""
response = client.chat.completions.create(
model="deepseek-chat",
messages=[
{"role": "system", "content": SYSTEM_PROMPT_TEXT},
{"role": "user", "content": combined_user_prompt}
],
temperature=0.1
)
markdown_result = response.choices[0].message.content.strip()
print(" ✅ Response received.")
except Exception as e:
print(f" ❌ API Error: {e}")
break
print(" [3/4] Validating output...")
# Validate against formatted text length, as it's closer to the expected output
is_valid, msg = validate_length(formatted_text, markdown_result)
if not is_valid:
print(f" 🛑 HALTED: {msg}")
break
print(f" ✅ {msg}")
print(" [4/4] Saving results...")
with open(md_filename, "a", encoding="utf-8") as f:
f.write(markdown_result)
f.write("\n\n")
save_progress(current_end)
print(f" ✅ Saved and checkpointed: pages [{current_start}-{current_end}]")
current_start = current_end + 1
if all_images_data:
print("\n" + "=" * 60)
print(f"💾 Saving {len(all_images_data)} image table(s)")
print("=" * 60)
try:
with open(images_md_filename, "w", encoding="utf-8") as f:
f.write("# Extracted Text from Images/Figures\n\n")
f.write("Column 2 is intentionally left empty for translation.\n\n---\n\n")
for img_data in all_images_data:
f.write(f"## Page {img_data['page']}, Image {img_data['image_index'] + 1}\n\n")
f.write(img_data['extracted_text'])
f.write("\n\n---\n\n")
print("✅ Image tables saved successfully!")
except Exception as e:
print(filed to save images file: {e}")
else:
print("\n⚠️ No text-containing images were found in this run.")
print("=" * 60)
print("🎉 PIPELINE COMPLETED SUCCESSFULLY")
print("=" * 60)
if __name__ == "__main__":
# if os.path.exists(PROGRESS_FILE): os.remove(PROGRESS_FILE)
process_pdf()
# %%
Итог обработки оставшихся 500 страниц: