《Python + Selenium 自动登录抓取入门》
1 简介
Selenium 通过 WebDriver 驱动真实浏览器完成页面自动化,常用于接口冒烟、UI 回归与需要登录态的抓取任务。Python 绑定 + ChromeDriver 是最常见的组合。
2 环境安装
2.1 安装浏览器与驱动
Debian 安装 Google Chrome:
wget -q -O - https://dl-ssl.google.com/linux/linux_signing_key.pub | apt-key add -
echo "deb http://dl.google.com/linux/deb/ stable main" >> /etc/apt/sources.list
apt-get update
apt-get install -y google-chrome-stable
安装 chromedriver 与 selenium:
apt-get install -y chromedriver python-pip
pip3 install selenium
chromedriver 版本需与 Chrome 主版本一致,可从 http://chromedriver.chromium.org/ 下载匹配版本并放到
/usr/local/bin。
2.2 初始化无头浏览器
from selenium import webdriver
def init_driver(driver_path="/usr/local/bin/chromedriver"):
option = webdriver.ChromeOptions()
option.add_argument("headless") # 无头模式,服务器上无需桌面
option.add_argument("--no-sandbox")
driver = webdriver.Chrome(executable_path=driver_path, options=option)
driver.implicitly_wait(10)
return driver
3 定位元素与登录操作
Selenium 支持多种锚定方式:id、name、class、XPath、CSS 等。用 XPath 定位输入框并输入账号密码:
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
def login(driver, url, account, password):
driver.get(url)
# 定位用户名输入框(以 GitHub 登录页为例)
input_user = driver.find_element_by_xpath("//*[@id=\"login_field\"]")
input_user.clear()
input_user.send_keys(account)
driver.find_element_by_xpath("//*[@id=\"password\"]").send_keys(password)
# 点击登录按钮
driver.find_element_by_xpath("//*[@id=\"login\"]/form/div[3]/input[3]").click()
4 显式等待
网络慢时元素可能尚未渲染,用显式等待代替 sleep:
def click_profile_menu(driver):
driver.find_element_by_xpath('//*[@id="user-links"]/li[3]/details/summary/img').click()
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located(
(By.XPATH, '//*[@id="user-links"]/li[3]/details/details-menu/ul/li[3]/a')
)).click()
5 下载页面资源
取元素的 src 属性并下载:
from urllib import request
def download_image(driver, css_selector, save_as="avatar.jpg"):
img = driver.find_element_by_css_selector(css_selector)
img_src = img.get_attribute("src")
request.urlretrieve(img_src, save_as)
6 完整脚本
#!/usr/bin/python3
# -*- coding: utf-8 -*-
import sys
from urllib import request
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
URL = "https://example.com/login"
ACCOUNT = sys.argv[1] if len(sys.argv) > 1 else ""
PASSWORD = sys.argv[2] if len(sys.argv) > 2 else ""
def init_driver(path):
option = webdriver.ChromeOptions()
option.add_argument("headless")
return webdriver.Chrome(executable_path=path, options=option)
def do_job(driver):
driver.get(URL)
driver.find_element_by_xpath("//input[@id='login_field']").send_keys(ACCOUNT)
driver.find_element_by_xpath("//input[@id='password']").send_keys(PASSWORD)
driver.find_element_by_xpath("//input[@type='submit']").click()
wait = WebDriverWait(driver, 10)
img = wait.until(EC.presence_of_element_located(
(By.CSS_SELECTOR, "img.avatar"))).get_attribute("src")
request.urlretrieve(img, "avatar.jpg")
if __name__ == "__main__":
driver = init_driver("/usr/local/bin/chromedriver")
try:
do_job(driver)
finally:
driver.quit()
7 常见问题
7.1 元素可定位但点击无效
- 确认页面有多个同名元素,改用更精确的 XPath / CSS 或
find_elements取目标; - 元素被遮挡时先
driver.execute_script("arguments[0].click();", el)。
7.2 找不到 chromedriver
- 将 chromedriver 放到 PATH 或显式传入
executable_path; - 确保其与本地 Chrome 主版本一致,否则启动即报
session not created。
8 学习资料
阅读 —
·
全站 —