I need a regexp I can use with PHP's preg_match_all() to match out content inside div-tags. The divs look like this:
<div id="t1">Content</div>
I've come up with this regexp so far which matches out all divs with id="t[number]"
/<div id="t(\\d)">(.*?)<\\/div>/
The problem is when the content consists of more divs, nested divs like this:
<div id="t1">Content <div>more stuff</div></div>
Any ideas on how I make my regexp work with nested tags?
Thanks
i think it will be better to use some DOM-instruments
As I recently found out, regex can't do that.
Matching pair tag with regex
I ended up using xpath, and it works like a charm
Try a parser instead:
Output:
Download the parser here: http://simplehtmldom.sourceforge.net/
Edit: More for my own amusement I tried to do it in regex. Here's what I came up with:
Output:
And a small explanation:
Now perhaps you understand why people try to persuade you from not using regex for this. As already noted, it will not help if the the html is improperly formed: the regex will make a bigger mess of the output than an html parser, I assure you. Also, the regex will probably make your eyes bleed and your colleagues (or the people who will maintain your software) may come looking for you after seeing what you did. :)
Your best bet is to first clean up your input (using TIDY or similar), and then use a parser to get the info you want.
If you believe this guy, there's at least one regex that does the trick, and he says it's faster than dom methods... I agree with him.
http://www.php.net/manual/fr/regexp.reference.recursive.php#95568