Productivity & Collaboration paper-to-slideseditable-pptxscientific-diagramsdrawio-exportpdf-conversionpresentation-generationknowledge-baseacademic-workflows

Paper2Any Research Workspace

Turn papers, text, and images into editable research figures, Draw.io diagrams, and slide decks.

FollowAgents review · FARS-2.1
Use with care
63/ 100 5-point scale 3.2 / 5
1 2 3 4 5 6
1Trust13 / 29 · 2.2/5

The documentation identifies external text, image, OCR, MinerU, Supabase, SAM3, and ONLYOFFICE services, persistent output mounts, and the distinction between the open-source repository and the hosted commercial implementation. This provides useful data-flow, effect, and provenance visibility. Deductions apply because papers, images, and API keys may be sent to several configurable services without retention, redaction, privacy-boundary, or provider-processing guidance; sharing BACKEND_API_KEY through VITE_API_KEY places that credential in a client build. No fine-grained permission model, security audit, dependency-vulnerability process, or operational rollback is shown. Script confirmation, outline editing, and deck version history provide only limited confirmation and recovery. Unknown publisher identity is not treated as suspicious.

2Reliability8 / 14 · 2.9/5

The manual smoke suite checks service health, HTTP failures, response shape, empty artifacts, and missing configuration with reasonably specific messages, while installation documentation identifies many service dependencies. Deductions apply because these tests are archived and manual, depend on a live backend and external models, and sometimes assert only that a path is a string or a value is nonempty. The README requires Python 3.11+, while pyproject permits 3.9+, and the Paper2Any product identity does not align cleanly with the dataflow-agent package name and description.

3Adaptability14 / 18 · 3.9/5

The evidence thoroughly identifies scenarios involving papers, text, topics, PDFs, images, and slides, with outputs spanning research figures, roadmaps, decks, posters, video, citations, and knowledge bases. Language, style, complexity, resolution, model, and provider choices are configurable. Deductions apply because some capabilities require SAM3, MinerU, OCR, TTS, ONLYOFFICE, or a commercial hosted layer, with boundaries disclosed but scattered. Triggering is largely explicit through UI choices and API parameters, without systematic suitability checks, conflict resolution, or input-risk constraints.

4Convention12 / 18 · 3.3/5

The README has a strong navigational structure with feature sections, showcases, quick start, project-structure and roadmap links, and contribution guidance. Docker and Linux instructions distinguish base, paper, CUDA, and operating-system dependencies, and the full Apache-2.0 license is present. Deductions apply because the supplied README ends partway through environment setup and provides limited FAQ evidence; Paper2Any, dataflow-agent, and the dfa command create naming drift. Dated news serves as a useful change summary but does not demonstrate a formal release, migration, or compatibility policy. An author email and contribution link only partially establish maintenance ownership, while publisher identity remains unknown.

5Effectiveness12 / 13 · 4.6/5

Editable PPTX, SVG, DrawIO XML, PNG, video, and batch-download outputs are directly reusable in editing and delivery workflows. The breadth of conversions offers substantial marginal value over manually reconstructing scientific visuals and presentations. Deductions apply because claims about one-click operation, layout accuracy, and speed are primarily showcased rather than supported by static quality benchmarks. Model APIs, GPU segmentation, Office infrastructure, and numerous system tools impose monetary, setup, and operational costs.

6Verifiability4 / 8 · 2.5/5

Several claims are traceable to named API endpoints, configuration variables, output fields, screenshots, and manual smoke cases. The README also clearly separates the repository from the not-fully-open Studio/Nexus product. Deductions apply because many visual-quality and performance claims lack linked implementation evidence, automated results, or benchmarks in the supplied material. The available tests are archived manual checks with some weak assertions, and package naming, description, and Python-version discrepancies limit cross-source corroboration.

Evidence confidence: Low Reviewed Aug 16, 2026 Reviewed revision 1ca1f658e811
The upstream repository has new commits since this review. The score still applies to the reviewed revision shown and may not cover the latest changes.
Before you use it
  • Before processing unpublished papers or sensitive material, verify retention, training-use, and logging policies for every model, OCR, MinerU, TTS, and hosted provider.
  • Do not treat VITE_API_KEY as a secret because frontend build variables are generally client-readable. Production deployments should use user-scoped authentication, short-lived tokens, server-side proxying, and key rotation.
  • Independently harden production deployments before using the example ONLYOFFICE settings that disable JWT, permit private-address access, or expose document download endpoints.
  • The hosted Studio/Nexus experience is not the same implementation as this repository; do not attribute hosted capabilities automatically to the open-source version.
  • Reconcile Python requirements, system tools, GPU services, and dependency compatibility before installation, and perform an independent dependency-vulnerability scan.
Review evidence [1][2][3][4][5][6]
See the full review method →

What does this agent do, and when should you use it?

Paper2Any is an open-source workspace for producing multimodal research assets, combining the dataflow_agent core, a FastAPI backend, a React frontend, command-line scripts, and tests. It accepts paper PDFs, long documents, text topics, images, and screenshots, then produces model diagrams, technical roadmaps, experimental plots, editable presentations, academic posters, video scripts, and technical reports. Distinct workflows—including Paper2Figure, Paper2Diagram, Paper2PPT, PDF2PPT, Image2PPT, PPTPolish, and the knowledge base—target different transformations and can emit PPTX, SVG, Draw.io, and PNG outputs. Users can work through the browser or invoke standalone CLI scripts, while model endpoints, keys, and model names are configurable for OpenAI or compatible API gateways. The stack supports Docker self-hosting as well as Linux, WSL, and native Windows installation, although document conversion and segmentation features require additional system tools or GPU-backed services. The hosted Paper to Any Studio / Nexus experience is separate, and its complete commercial implementation is not included in this repository.

Paper2Any reads PDFs, images, screenshots, free-form text, or topics and routes them through the agentroles, workflow, promptstemplates, and toolkits modules under dataflow_agent. Paper2Figure generates model_arch, tech_route, or exp_data figures; Paper2Diagram / Image2Drawio converts papers, text, or images into chat-editable Draw.io canvases with drawio, PNG, and SVG export. Paper2PPT handles papers, long documents, and topics, extracts tables and figures, assists with outline editing, and builds editable decks; PDF2PPT preserves source layouts, while Image2PPT reconstructs images as structured slides. PPTPolish applies layout optimization and style transfer to existing presentations, while Paper2Poster, Paper2Video, Paper2Rebuttal, Paper2Citation, and Paper2Technical cover posters, narration scripts, review responses, citation exploration, and technical summaries. The knowledge-base path performs ingestion, embedding, and semantic search before generating PPT, podcast, or mind-map material. In web deployments, frontend-workflow calls fastapi_app through the /api/v1/ API; CLI users can run script/run_paper2figure_cli.py, script/run_paper2ppt_cli.py, script/run_pdf2ppt_cli.py, script/run_image2ppt_cli.py, and script/run_ppt2polish_cli.py directly.

  1. A researcher preparing a lab meeting can convert a paper PDF into an editable deck with extracted tables and figures, then refine slide content on the canvas.
  2. A paper author who needs a method illustration can use Paper2Figure to generate a model architecture, technical roadmap, or experimental-data plot from the manuscript or a text description.
  3. An engineering team reconstructing an existing diagram can send a screenshot through Image2Drawio, receive an editable Draw.io structure, and revise it through chat.
  4. A laboratory maintaining an internal document collection can ingest and embed files, run semantic searches, and generate knowledge-base-grounded presentations, podcasts, or mind maps.
  5. An author preparing submission materials can produce an editable academic poster or use Paper2Rebuttal to draft responses grounded in claims and evidence.
  6. An instructor or technical communicator modernizing old material can use PDF2PPT, Image2PPT, or PPTPolish to recover layouts and apply a consistent visual style.

What are this agent's strengths and limitations?

Pros
  • Editability is a first-class output requirement: research figures support PPT and SVG, Draw.io workflows export drawio, PNG, and SVG, and presentation workflows produce editable PPTX files.
  • One repository covers papers, long documents, text topics, PDFs, images, and screenshots through both browser and CLI interfaces.
  • Model configuration is not tied to one hard-coded model: deployments can select API URLs, credentials, and models through unified or workflow-specific settings.
  • The knowledge base goes beyond ingestion and semantic search by feeding retrieved material into presentation, podcast, and mind-map generation.
  • Docker scripts cover building, startup, logs, shutdown, and persistent output and model directories.
Limitations
  • A complete installation has substantial native dependencies, including LibreOffice, Inkscape, ffmpeg, Poppler, wkhtmltopdf, and Tectonic; installing the Python package alone does not enable every feature.
  • PDF2PPT, Image2PPT, and Image2Drawio depend on a SAM3 segmentation service, which may require NVIDIA GPUs, downloaded model assets, and additional service configuration.
  • Most generation workflows rely on external text or image APIs, creating dependencies on network access, credentials, provider availability, and potentially metered usage.
  • Native Windows is not the preferred deployment route; Linux or WSL is recommended, and Windows acceleration requires a separately matched vLLM wheel.
  • The hosted Studio / Nexus product is not fully represented in the repository, so self-hosters should not expect its Spaces, Labs, Credits, or desktop-release experience.
  • The published roadmap marks several areas as incomplete: Paper2PPT is listed at 70%, PPTPolish at 60%, and Paper2Video at 40%, with some optimization and asset features still in progress.

How do you install or deploy this agent?

The documented self-hosted path uses Docker:

  1. Run git clone https://github.com/OpenDCAI/Paper2Any.git, followed by cd Paper2Any.
  2. Run cp fastapi_app/.env.simple.example fastapi_app/.env, cp frontend-workflow/.env.simple.example frontend-workflow/.env, and cp deploy/docker.env.example deploy/docker.env.
  3. Set BACKEND_API_KEY, SIMPLE_TEXT_API_URL, and SIMPLE_TEXT_API_KEY in the backend environment file. Add SIMPLE_IMAGE_API_URL and SIMPLE_IMAGE_API_KEY when image generation is required.
  4. Set frontend VITE_API_KEY to the same value as BACKEND_API_KEY.
  5. Run bash deploy/docker-up.sh, open http://localhost:3000, and verify the backend at http://localhost:8000/health.

For a manual Linux installation, run conda create -n paper2any python=3.11 -y, conda activate paper2any, pip install -r requirements-base.txt, pip install -e ., and pip install -r requirements-paper.txt. Ubuntu or Debian hosts also need ffmpeg inkscape libreoffice poppler-utils wkhtmltopdf, with Tectonic recommended through Conda. PDF2PPT, Image2PPT, and Image2Drawio additionally require an external SAM3_SERVER_URLS endpoint or the optional local service started with DOCKER_WITH_SAM3=1 bash deploy/docker-up.sh.

How do you use this agent?

For the shortest CLI path, set export DF_API_URL=https://api.openai.com/v1, export DF_API_KEY=sk-xxx, and export DF_MODEL=gpt-4o. Create a 15-slide paper presentation with python script/run_paper2ppt_cli.py --input paper.pdf --api-key sk-xxx --page-count 15. Generate a model architecture with python script/run_paper2figure_cli.py --input paper.pdf --graph-type model_arch --api-key sk-xxx. Convert a PDF into an editable presentation with python script/run_pdf2ppt_cli.py --input slides.pdf; append --use-ai-edit --api-key sk-xxx for AI enhancement. Browser users can start the stack, visit http://localhost:3000, select the relevant Paper2Figure, Paper2PPT, Drawio, or conversion workflow, and upload the source material. Ensure SAM3 is reachable before using the three segmentation-dependent workflows. ONLYOFFICE browser editing is optional and requires a separately deployed ONLYOFFICE Document Server plus the documented backend variables.

How does this agent compare with similar options?

Compared with the hosted Paper to Any Studio / Nexus, this repository supplies the self-hostable open-source Paper2Any code, web stack, and CLI tools, but not the complete commercial Studio implementation. Within Paper2Any, simple configuration provides unified text and image endpoints, whereas advanced .env.example files allow workflow-specific model and provider overrides at the cost of more configuration.

FAQ

Is Supabase required for the core workflows?
No. Core features and CLI scripts work without Supabase, but authentication, points, redemption, invitations, history, and cloud file storage are unavailable.
Does every workflow require a GPU?
No. Paper2PPT, Paper2Figure, and knowledge-base workflows can operate through remote LLM APIs. PDF2PPT, Image2PPT, and Image2Drawio require SAM3, supplied either as an external service or an optional local GPU container.
Will model usage incur charges?
The repository does not define a single price. Most generation tasks require a user-supplied text or image model API, so costs depend on the selected provider, model, and usage volume.
Are the generated files actually editable?
The main workflows expose editable formats: Paper2Figure can output PPT and SVG, Paper2Diagram exports Draw.io files, and Paper2PPT, PDF2PPT, and Image2PPT produce editable PPTX. ONLYOFFICE online editing is an optional deployment layer.
Is the online demo identical to the self-hosted repository?
No. The public demo leads to Paper to Any Studio / Nexus, while this repository remains the open-source Paper2Any codebase and does not include the full commercial Studio implementation.

Related agents