Generate bigrams with NLTK

2019-03-25 14:55发布

I am trying to produce a bigram list of a given sentence for example, if I type,

    To be or not to be

I want the program to generate

     to be, be or, or not, not to, to be

I tried the following code but just gives me

<generator object bigrams at 0x0000000009231360>

This is my code:

    import nltk
    bigrm = nltk.bigrams(text)
    print(bigrm)

So how do I get what I want? I want a list of combinations of the words like above (to be, be or, or not, not to, to be).

标签： python nltk n-gram

2条回答

ゆ、 Hurt°

2楼-- · 2019-03-25 15:18

The following code produce a bigram list for a given sentence

>>> import nltk
>>> from nltk.tokenize import word_tokenize
>>> text = "to be or not to be"
>>> tokens = nltk.word_tokenize(text)
>>> bigrm = nltk.bigrams(tokens)
>>> print(*map(' '.join, bigrm), sep=', ')
to be, be or, or not, not to, to be

0人赞添加讨论(0) 举报

闹够了就滚

3楼-- · 2019-03-25 15:21

nltk.bigrams() returns an iterator (a generator specifically) of bigrams. If you want a list, pass the iterator to list(). It also expects a sequence of items to generate bigrams from, so you have to split the text before passing it (if you had not done it):

bigrm = list(nltk.bigrams(text.split()))

To print them out separated with commas, you could (in python 3):

print(*map(' '.join, bigrm), sep=', ')

If on python 2, then for example:

print ', '.join(' '.join((a, b)) for a, b in bigrm)

Note that just for printing you do not need to generate a list, just use the iterator.

0人赞添加讨论(0) 举报

Generate bigrams with NLTK

采纳回答

编辑标签

举报内容

检举类型

检举原因

检举说明(必填)

打开微信“扫一扫”，打开网页后点击屏幕右上角分享按钮

付费偷看金额在0.1-10元之间