你可以從下面的例子確切的結果。
使用python next兄弟來獲得合適的結果。
的HTML代碼是:
<div id="provider-region-addresses">
<h3>Contact details</h3>
<h2 class="toggler nohide">Auckland</h2>
<dl class="clear">
<dt>More information</dt>
<dd>North Shore Hospital</dd><dt>Physical address</dt>
<dd>124 Shakespeare Rd, Takapuna, Auckland 0620</dd><dt>Postal address</dt>
<dd>Private Bag 93503, Takapuna, Auckland 0740</dd><dt>Postcode</dt>
<dd>0740</dd><dt>District/town</dt>
<dd>
North Shore, Takapuna</dd><dt>Region</dt>
<dd>Auckland</dd><dt>Phone</dt>
<dd>(09) 486 8996</dd><dt>Fax</dt>
<dd>(09) 486 8342</dd><dt>Website</dt>
<dd><a target="_blank" href="http://www.healthpoint.co.nz/default,61031.sm">http://www.healthpoint.co.nz/default,61031...</a></dd>
</dl>
<h2 class="toggler nohide">Auckland</h2>
<dl class="clear">
<dt>Physical address</dt>
<dd>Helensville</dd><dt>Postal address</dt>
<dd>PO Box 13, Helensville 0840</dd><dt>Postcode</dt>
<dd>0840</dd><dt>District/town</dt>
<dd>
Rodney, Helensville</dd><dt>Region</dt>
<dd>Auckland</dd><dt>Phone</dt>
<dd>(09) 420 9450</dd><dt>Fax</dt>
<dd>(09) 420 7050</dd><dt>Website</dt>
<dd><a target="_blank" href="http://www.healthpoint.co.nz/default,61031.sm">http://www.healthpoint.co.nz/default,61031...</a></dd>
</dl>
<h2 class="toggler nohide">Auckland</h2>
<dl class="clear">
<dt>Physical address</dt>
<dd>Warkworth</dd><dt>Postal address</dt>
<dd>PO Box 505, Warkworth 0941</dd><dt>Postcode</dt>
<dd>0941</dd><dt>District/town</dt>
<dd>
Rodney, Warkworth</dd><dt>Region</dt>
<dd>Auckland</dd><dt>Phone</dt>
<dd>(09) 422 2700</dd><dt>Fax</dt>
<dd>(09) 422 2709</dd><dt>Website</dt>
<dd><a target="_blank" href="http://www.healthpoint.co.nz/default,61031.sm">http://www.healthpoint.co.nz/default,61031...</a></dd>
</dl>
<h2 class="toggler nohide">Auckland</h2>
<dl class="clear">
<dt>More information</dt>
<dd>Waitakere Hospital</dd><dt>Physical address</dt>
<dd>55-75 Lincoln Rd, Henderson, Auckland 0610</dd><dt>Postal address</dt>
<dd>Private Bag 93115, Henderson, Auckland 0650</dd><dt>Postcode</dt>
<dd>0650</dd><dt>District/town</dt>
<dd>
Waitakere, Henderson</dd><dt>Region</dt>
<dd>Auckland</dd><dt>Phone</dt>
<dd>(09) 839 0000</dd><dt>Fax</dt>
<dd>(09) 837 6634</dd><dt>Website</dt>
<dd><a target="_blank" href="http://www.healthpoint.co.nz/default,61031.sm">http://www.healthpoint.co.nz/default,61031...</a></dd>
</dl>
<h2 class="toggler nohide">Auckland</h2>
<dl class="clear">
<dt>More information</dt>
<dd>Hibiscus Coast Community Health Centre</dd><dt>Physical address</dt>
<dd>136 Whangaparaoa Rd, Red Beach 0932</dd><dt>Postcode</dt>
<dd>0932</dd><dt>District/town</dt>
<dd>
Rodney, Red Beach</dd><dt>Region</dt>
<dd>Auckland</dd><dt>Phone</dt>
<dd>(09) 427 0300</dd><dt>Fax</dt>
<dd>(09) 427 0391</dd><dt>Website</dt>
<dd><a target="_blank" href="http://www.healthpoint.co.nz/default,61031.sm">http://www.healthpoint.co.nz/default,61031...</a></dd>
</dl>
</div>
蜘蛛的代碼是:
def parse(self, response):
hxs = HtmlXPathSelector(response)
practice = hxs.select('//h1/text()').extract()
items1 = []
results = hxs.select('//*[@id="content"]/div[@class="content"]/div/dl')
for result in results:
item = WebhealthItem1()
#item['url'] = result.select('//dl/a/@href').extract()
item['practice'] = practice
item['hours'] = map(unicode.strip,
result.select('dt[contains(.," Contact hours")]/following-sibling::dd[1]/text()').extract())
item['more_hours'] = map(unicode.strip,
result.select('dt[contains(., "More information")]/following-sibling::dd[1]/text()').extract())
item['physical_address'] = map(unicode.strip,
result.select('dt[contains(., "Physical address")]/following-sibling::dd[1]/text()').extract())
item['postal_address'] = map(unicode.strip,
result.select('dt[contains(., "Postal address")]/following-sibling::dd[1]/text()').extract())
item['postcode'] = map(unicode.strip,
result.select('dt[contains(., "Postcode")]/following-sibling::dd[1]/text()').extract())
item['district_town'] = map(unicode.strip,
result.select('dt[contains(., "District/town")]/following-sibling::dd[1]/text()').extract())
item['region'] = map(unicode.strip,
result.select('dt[contains(., "Region")]/following-sibling::dd[1]/text()').extract())
item['phone'] = map(unicode.strip,
result.select('dt[contains(., "Phone")]/following-sibling::dd[1]/text()').extract())
item['website'] = map(unicode.strip,
result.select('dt[contains(., "Website")]/following-sibling::dd[1]/a/@href').extract())
item['email'] = map(unicode.strip,
result.select('dt[contains(., "Email")]/following-sibling::dd[1]/a/text()').extract())
items1.append(item)
return items1
這似乎是更喜歡它,但它並沒有拿起每一個號碼。它只顯示11條記錄?另外我應該如何發現關於battingComms類?謝謝 – Del 2014-09-30 17:28:41
@Del,我怎麼知道你想從頁面上得到什麼? – alecxe 2014-09-30 17:34:00
如果我不清楚,我很抱歉。我想要一個帶有2列的csv文件。一列說數字:0.1,0.3 ... 19.5,19.6。另一欄顯示網頁上該號碼旁邊的文字。 – Del 2014-09-30 18:57:29