I want to retrieve all percentage data as well as integer/float numbers with units from an input text, if it is present in the text. If both are not present together, I want to retrieve atleast the one that is present. Till now if there is an integer/float with an unit in the extracted text, it comes in the result variable.
result=[]
newregex = "[0-9\.\s]+(?:mg|kg|ml|q.s.|ui|M|g|µg)"
percentregex = "(\d+(\.\d+)?%)"
for s in zz:
for e in extracteddata:
v = re.search(newregex,e,flags=re.IGNORECASE|re.MULTILINE)
xx = re.search(percentregex,e,flags=re.IGNORECASE|re.MULTILINE)
if v:
if e.upper().startswith(s.upper()):
result.append([s,v.group(0), e])
else:
if e.upper().startswith(s.upper()):
result.append([s, e])
In the code above, newregex identifies numbers/float with an unit after it, percentregex identifies percentage data, zz and extracteddata are as follows
zz = ['HYDROCHLORIC ACID 2M', 'ROPIVACAINE HYDROCHLORIDE MONOHYDRATE', 'SODIUM CHLORIDE', 'SODIUM HYDROXIDE 2M', 'WATER FOR INJECTIONS']
extracteddata = ['Ropivacaine hydrochloride monohydrate for injection (corresponding to 2 mg Ropivacaine hydrochloride anhydrous) 2.12 mg Active ingredient Ph Eur ', 'Sodium chloride for injection 8.6 mg 28% Tonicity contributor Ph Eur ', 'Sodium hydroxide 2M q.s. pH-regulator Ph Eur, NF Hydrochloric acid 2M q.s. pH-regulator Ph Eur, NF ', 'Water for Injections to 1 ml 34% Solvent Ph Eur, USP The product is filled into polypropylene bags sealed with rubber stoppers and aluminium caps with flip-off seals. The primary container is enclosed in a blister. 1(1)']
Now I also want to add the condition to extract percentage data in the result variable if it is present but I am stuck with the looping aspect. i want help on using the variable 'xx' to add percentage data to result list if it is present, along with the integer/float numbers with units.
Any help on this.
Updates on attempts made:
result = []
mg = []
newregex = "[0-9\.\s]+(?:mg|kg|ml|q.s.|ui|M|g|µg)"
percentregex = "(\d+(\.\d+)?%)"
print(type(newregex))
for s in zz:
for e in extracteddata:
v = re.search(newregex,e,flags=re.IGNORECASE|re.MULTILINE)
xx = re.search(percentregex,e,flags=re.IGNORECASE|re.MULTILINE)
if v:
# mg.append(v.group(0))
if e.upper().startswith(s.upper()):
result.append([s,v.group(0), e])
elif v is None:
if e.upper().startswith(s.upper()):
result.append([s, e])
elif xx:
if v:
if e.upper().startswith(s.upper()):
result.append([s,v.group(0),xx.group(0), e])
elif v is None:
if xx:
if e.upper().startswith(s.upper()):
result.append([s,xx.group(0), e])
elif v is None and xx is None:
if e.upper().startswith(s.upper()):
result.append([s, e])
else:
print("DOne")
Here is a Python demo of what we talked about in the comments :
mod per request
This is the regex expanded
As seen it uses three look ahead assertions to find the first instances
of the unit and percentage numbers and stand alone numbers.
All values are unique and not an overlap.
Testing each one for non-empty shows if it found that item(s) in the line.