How to break up a paragraph by sentences in Python

2019-01-17 18:26发布

I need to parse sentences from a paragraph in Python. Is there an existing package to do this, or should I be trying to use regex here?

标签： python regex text-segmentation

2条回答

傲

2楼-- · 2019-01-17 18:44

The nltk.tokenize module is designed for this and handles edge cases. For example:

>>> from nltk import tokenize
>>> p = "Good morning Dr. Adams. The patient is waiting for you in room number 3."
>>> tokenize.sent_tokenize(p)
['Good morning Dr. Adams.', 'The patient is waiting for you in room number 3.']

0人赞添加讨论(0) 举报

beautiful°

3楼-- · 2019-01-17 19:02

Here is how I am getting the first n sentences:

def get_first_n_sentence(text, n):
    endsentence = ".?!"
    sentences = itertools.groupby(text, lambda x: any(x.endswith(punct) for punct in endsentence))
    for number,(truth, sentence) in enumerate(sentences):
        if truth:
            first_n_sentences = previous+''.join(sentence).replace('\n',' ')
        previous = ''.join(sentence)
        if number>=2*n: break #

    return first_n_sentences

Reference: http://www.daniweb.com/software-development/python/threads/303844

0人赞添加讨论(0) 举报

How to break up a paragraph by sentences in Python

采纳回答

编辑标签

举报内容

检举类型

检举原因

检举说明(必填)

打开微信“扫一扫”，打开网页后点击屏幕右上角分享按钮

付费偷看金额在0.1-10元之间