如何提取以电话开头并以 } 结尾的短语



如何使用正则表达式和 Python 提取以 phone 开头并以"}"结尾的短语

我试图从页面源中提取数据。

{"meta":{"subtitle":"Apartment for Rent in Marina Gate 1, Marina Gate","price":145000,"price_text":"145,000 AED/year","contact_options":{"list":{"phone":{"type":"phone","value":"+XXXXXXXX","link":"tel:+XXXXXXXX","is_did":true},"email":{"type":"email","value":"name@email.com","link":"mailto:name@email.com"}},"details":{"phone":{"type":"phone","value":"+XXXXXXXX","link":"tel:+XXXXXXXX","is_did":true},"sms":{"type":"sms","value":"+XXXXXXXX","link":"sms:+XXXXXXXX"},"email":{"type":"email","value":"name@email.com","link":"mailto:name@email.com"}}},"images_count":11}}'

我想用正则表达式提取所有以电话开头并以 } 结尾的短语

我试过 re.findall(r"^phone(.*(}$",source(

这就是我想要的"电话","值":"+XXXXXXXX","链接":"tel:+XXXXXXXX","is_did":true}

为此

使用 json 而不是正则表达式可能更好。试试这个,

import json
test_str = '{"meta":{"subtitle":"Apartment for Rent in Marina Gate 1, Marina Gate","price":145000,"price_text":"145,000 AED/year","contact_options":{"list":{"phone":{"type":"phone","value":"+XXXXXXXX","link":"tel:+XXXXXXXX","is_did":true},"email":{"type":"email","value":"name@email.com","link":"mailto:name@email.com"}},"details":{"phone":{"type":"phone","value":"+XXXXXXXX","link":"tel:+XXXXXXXX","is_did":true},"sms":{"type":"sms","value":"+XXXXXXXX","link":"sms:+XXXXXXXX"},"email":{"type":"email","value":"name@email.com","link":"mailto:name@email.com"}}},"images_count":11}}'
print test_str
json_str = json.loads(test_str)
print json_str
phone_num = json_str['meta']['contact_options']['list']['phone']
print phone_num

你可以试试这段代码(用re解析<script>标签(:

import requests
import json
import re
html_text = requests.get('https://www.propertyfinder.ae/en/rent/apartment-for-rent-dubai-dubai-marina-marina-gate-marina-gate-1-6951117.html').text
data = json.loads(re.findall(r'payloads*:s*(.*?)n', html_text)[0])
for d in data['data']:
    print(d['meta']['subtitle'])
    print(d['meta']['contact_options'])
    print('*' * 80)

指纹:

Apartment for Rent in Marina Gate 1, Marina Gate
{'list': {'phone': {'type': 'phone', 'value': '+971528347286', 'link': 'tel:+971528347286', 'is_did': True}, 'email': {'type': 'email', 'value': 'ahmad@providentestate.com', 'link': 'mailto:ahmad@providentestate.com'}}, 'details': {'phone': {'type': 'phone', 'value': '+971528347286', 'link': 'tel:+971528347286', 'is_did': True}, 'sms': {'type': 'sms', 'value': '+971581806000', 'link': 'sms:+971581806000'}, 'email': {'type': 'email', 'value': 'ahmad@providentestate.com', 'link': 'mailto:ahmad@providentestate.com'}}}
********************************************************************************
Apartment for Rent in Marina Gate 1, Marina Gate
{'list': {'phone': {'type': 'phone', 'value': '+971525226138', 'link': 'tel:+971525226138', 'is_did': True}, 'email': {'type': 'email', 'value': 'suhail.p@w2realestate.com', 'link': 'mailto:suhail.p@w2realestate.com'}}, 'details': {'phone': {'type': 'phone', 'value': '+971525226138', 'link': 'tel:+971525226138', 'is_did': True}, 'sms': {'type': 'sms', 'value': '+971503940533', 'link': 'sms:+971503940533'}, 'email': {'type': 'email', 'value': 'suhail.p@w2realestate.com', 'link': 'mailto:suhail.p@w2realestate.com'}}}
********************************************************************************
Apartment for Rent in Marina Gate 1, Marina Gate
{'list': {'phone': {'type': 'phone', 'value': '+971528347286', 'link': 'tel:+971528347286', 'is_did': True}, 'email': {'type': 'email', 'value': 'ahmad@providentestate.com', 'link': 'mailto:ahmad@providentestate.com'}}, 'details': {'phone': {'type': 'phone', 'value': '+971528347286', 'link': 'tel:+971528347286', 'is_did': True}, 'sms': {'type': 'sms', 'value': '+971581806000', 'link': 'sms:+971581806000'}, 'email': {'type': 'email', 'value': 'ahmad@providentestate.com', 'link': 'mailto:ahmad@providentestate.com'}}}
********************************************************************************
Apartment for Rent in Marina Gate 1, Marina Gate
{'list': {'phone': {'type': 'phone', 'value': '+971522233791', 'link': 'tel:+971522233791', 'is_did': True}, 'email': {'type': 'email', 'value': 'eddy@exclusive-links.com', 'link': 'mailto:eddy@exclusive-links.com'}}, 'details': {'phone': {'type': 'phone', 'value': '+971522233791', 'link': 'tel:+971522233791', 'is_did': True}, 'sms': {'type': 'sms', 'value': '+971523279984', 'link': 'sms:+971523279984'}, 'email': {'type': 'email', 'value': 'eddy@exclusive-links.com', 'link': 'mailto:eddy@exclusive-links.com'}}}
********************************************************************************
Apartment for Rent in Marina Gate 1, Marina Gate
{'list': {'phone': {'type': 'phone', 'value': '+971565775168', 'link': 'tel:+971565775168', 'is_did': False}, 'email': {'type': 'email', 'value': 'julia@abodeproperty.ae', 'link': 'mailto:julia@abodeproperty.ae'}}, 'details': {'phone': {'type': 'phone', 'value': '+971565775168', 'link': 'tel:+971565775168', 'is_did': False}, 'sms': {'type': 'sms', 'value': '+971565775168', 'link': 'sms:+971565775168'}, 'email': {'type': 'email', 'value': 'julia@abodeproperty.ae', 'link': 'mailto:julia@abodeproperty.ae'}}}
********************************************************************************

注意:有时网站会返回格式错误的 HTML 代码,因此您需要多次运行脚本直到成功(也许您需要调整regex - 我没有进一步调查(

最新更新