2016-11-16 57 views
0

我試圖用slate模塊來提取PDF文件中的文本提取PDF文本,如本使用python3

$sudo pip install https://codeload.github.com/timClicks/slate/zip/master 
Collecting https://codeload.github.com/timClicks/slate/zip/master 
    Downloading https://codeload.github.com/timClicks/slate/zip/master 
Requirement already satisfied: distribute in /usr/lib/python3.5/site-packages (from slate==0.5.2) 
Requirement already satisfied: pdfminer3k in /usr/lib/python3.5/site-packages (from slate==0.5.2) 
Requirement already satisfied: setuptools>=0.7 in /usr/lib/python3.5/site-packages (from distribute->slate==0.5.2) 
Requirement already satisfied: pytest>=2.0 in /usr/lib/python3.5/site-packages (from pdfminer3k->slate==0.5.2) 
Requirement already satisfied: ply>=3.4 in /usr/lib/python3.5/site-packages (from pdfminer3k->slate==0.5.2) 
Requirement already satisfied: py>=1.4.29 in /usr/lib/python3.5/site-packages (from pytest>=2.0->pdfminer3k->slate==0.5.2) 
Installing collected packages: slate 
    Found existing installation: slate 0.3 
    Uninstalling slate-0.3: 
     Successfully uninstalled slate-0.3 
    Running setup.py install for slate ... done 
Successfully installed slate-0.5.2 

,我想:

#!/usr/bin/python3 
import slate 

with open('/var/tmp/PhysRevB.93.014203.pdf') as fp: 
    doc = slate.PDF(fp) 
print(len(doc)) 
print(doc[0]) 

這是給我的錯誤:

$python3 tstslt.py 
Traceback (most recent call last): 
    File "tstslt.py", line 2, in <module> 
    import slate 
    File "/usr/lib/python3.5/site-packages/slate/__init__.py", line 66, in <module> 
    from .classes import PDF 
    File "/usr/lib/python3.5/site-packages/slate/classes.py", line 25, in <module> 
    import utils 
ImportError: No module named 'utils' 

我可以使用PyPDF2來提取文本,但看看是否s晚點比較好。

回答

0

根據this issue一個平板的依賴條件的(pdfminer)不支持Python3

(...)

The "pdfminer" that is required does not work because it is currently incompatible with python 3.5.

It says so on their readme:

https://github.com/euske/pdfminer

"Install Python 2.6 or newer. (Python 3 is not supported.)"

+0

雖然這種聯繫可以回答這個問題,最好是在這裏有答案的主要部件,並提供鏈接參考。如果鏈接頁面更改,則僅鏈接答案可能會失效。 - [來自評論](/ review/low-quality-posts/17579044) –