Python regex for finding contents of MediaWiki mar

If I have some xml containing things like the following mediawiki markup:

" ...collected in the 12th century, of which [[Alexander the Great]] was the hero, and in which he was represented, somewhat like the British [[King Arthur|Arthur]]"

what would be the appropriate arguments to something like:

re.findall([[__?__]], article_entry)

I am stumbling a bit on escaping the double square brackets, and getting the proper link for text like: [[Alexander of Paris|poet named Alexander]]

标签： python regex mediawiki

4条回答

我命由我不由天

2楼-- · 2020-07-22 10:46

RegExp: \w+( \w+)+(?=]])

input

[[Alexander of Paris|poet named Alexander]]

output

poet named Alexander

input

[[Alexander of Paris]]

output

Alexander of Paris

0人赞添加讨论(0) 举报

闹够了就滚

3楼-- · 2020-07-22 10:49

If you are trying to get all the links from a page, of course it is much easier to use the MediaWiki API if at all possible, e.g. http://en.wikipedia.org/w/api.php?action=query&prop=links&titles=Stack_Overflow_(website).

Note that both these methods miss links embedded in templates.

0人赞添加讨论(0) 举报

不美不萌又怎样

4楼-- · 2020-07-22 10:54

Here is an example

import re

pattern = re.compile(r"\[\[([\w \|]+)\]\]")
text = "blah blah [[Alexander of Paris|poet named Alexander]] bldfkas"
results = pattern.findall(text)

output = []
for link in results:
    output.append(link.split("|")[0])

# outputs ['Alexander of Paris']

Version 2, puts more into the regex, but as a result, changes the output:

import re

pattern = re.compile(r"\[\[([\w ]+)(\|[\w ]+)?\]\]")
text = "[[a|b]] fdkjf [[c|d]] fjdsj [[efg]]"
results = pattern.findall(text)

# outputs [('a', '|b'), ('c', '|d'), ('efg', '')]

print [link[0] for link in results]

# outputs ['a', 'c', 'efg']

Version 3, if you only want the link without the title.

pattern = re.compile(r"\[\[([\w ]+)(?:\|[\w ]+)?\]\]")
text = "[[a|b]] fdkjf [[c|d]] fjdsj [[efg]]"
results = pattern.findall(text)

# outputs ['a', 'c', 'efg']

0人赞添加讨论(0) 举报

家丑人穷心不美

5楼-- · 2020-07-22 10:54

import re
pattern = re.compile(r"\[\[([\w ]+)(?:\||\]\])")
text = "of which [[Alexander the Great]] was somewhat like [[King Arthur|Arthur]]"
results = pattern.findall(text)
print results

Would give the output

["Alexander the Great", "King Arthur"]

0人赞添加讨论(0) 举报

Python regex for finding contents of MediaWiki mar

采纳回答

编辑标签

举报内容

检举类型

检举原因

检举说明(必填)

打开微信“扫一扫”，打开网页后点击屏幕右上角分享按钮

付费偷看金额在0.1-10元之间