BeautifulSoup を使用してスクレイピングされたテキストを Pandas データフレームに変換する

debugcn 投稿 Dev

プラシャーント・マノハル

以下のコードを使用して、Web サイトからテキストを抽出しています。私はそれを文字列の形で持っています。

import requests
URL = 'https://www.instituteforsupplymanagement.org/about/MediaRoom/newsreleasedetail.cfm?ItemNumber=30655&SSO=1'
r = requests.get(URL)
page = r.text

from bs4 import BeautifulSoup
soup = BeautifulSoup(page, 'lxml')
import re

strong_el = soup.find('strong',text='WHAT RESPONDENTS ARE SAYING …')

ul_tag = strong_el.find_next_sibling('ul')
LI_TAG =''
for li_tag in ul_tag.children:

    LI_TAG += li_tag.string

print LI_TAG

2 列のデータフレームを作成しようとしています: 1) コメント 2) 業界 (括弧内のサブ文字列)。次のようにStringIOを使用しようとしたときにエラーが発生しました: 「TypeError: data argument can't be an iterator」。これらのコメントをデータフレームに変換するにはどうすればよいですか?

import sys
if sys.version_info[0] < 3: 
    from StringIO import StringIO
else:
    from io import StringIO

import pandas as pd

LI_TAG = StringIO(LI_TAG)
df = pd.DataFrame(LI_TAG)

ロビー

LI_TAG 変数は単なる長い文字列のようです - したがって、データフレームに格納するには分割する必要があります。

import requests
URL = 'https://www.instituteforsupplymanagement.org/about/MediaRoom/newsreleasedetail.cfm?ItemNumber=30655&SSO=1'
r = requests.get(URL)
page = r.text

from bs4 import BeautifulSoup
soup = BeautifulSoup(page, 'lxml')
import re

strong_el = soup.find('strong',text='WHAT RESPONDENTS ARE SAYING …')

ul_tag = strong_el.find_next_sibling('ul')
LI_TAG =''
for li_tag in ul_tag.children:

    LI_TAG += li_tag.string

# Convert to unicode to remove quotation marks \u201c and \u201d
LI_TAG_U = unicode(LI_TAG)
comments=[]
industries=[]
for string in LI_TAG.strip().split('\n'):
    comment, industry =  string.split(u'\u201d')
    comments.append(comment.strip(u'\u201c'))
    industries.append(industry.strip(' (').strip(')'))

import pandas as pd

data = pd.DataFrame()

data['Comment']=comments
data['Industry']=industries

これがあなたのために働くことを願っています!

この記事はインターネットから収集されたものであり、転載の際にはソースを示してください。

侵害の場合は、連絡してください[email protected]