Skip to content

Docsplit::TextExtractor#extract_text should return the path of the output text file? #139

Description

@nruth

related to #42

After extracting the text from a PDF or Doc file I need to do something with it. I understand not loading the string into ruby (it could be huge), but it'd be helpful to get the output file path as a return value. Otherwise we have to use different output dirs or try to reconstruct its path based on other information, which feels wrong.

Currently Docsplit::TextExtractor#extract_text is returning the source file paths. For Transparent doc(x) file conversion it returns the intermediary tempfile pdf.
E.g. when I map over an array with a pdf and a doc in my project's tmp dir I get back

[
"/var/folders/_j/q3pr8b3s1vj85mhqvyb06gr40000gn/T/docsplit/sample.docx20160125-29577-go3upi.pdf",
"/Users/nruth/dev/monitor/tmp/AISB08.pdf20160125-29577-1svhpfo.pdf"
]

Instead I'd like to be given the path of the output text files, so I can open them.

Would this be a good PR, or is there a deliberate reason to return these other file paths that could be documented?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions