如何用Java提取网页表格信息并生成自定义类对象用于Android Studio
可行解决方案
方案一:爬虫提取+JSON解析生成列表
1. 用Python提取网页表格数据
用BeautifulSoup解析目标网页的物品表格,对应字段映射:
- Name:表格中的物品名称单元格文本
- Id:从表格行的属性或物品详情页URL中提取(比如
?id=xxx参数) - Icon:提取物品图片的
src链接 - Quote:物品对应的引用文本(通常在表格行的特定单元格)
- Description:物品的效果描述文本
- Quality:提取品质数值(注意转成int类型)
提取完成后,将数据导出为items.json文件,格式示例:
[ { "name": "Sad Onion", "id": 1, "iconUrl": "https://static.wikia.nocookie.net/bindingofisaacre_gamepedia/images/7/7e/Sad_Onion.png", "quote": "\"I can cry more tears!\"", "description": "Increases tear rate by 33%.", "quality": 1 } ]
2. Android端解析生成对象列表
首先创建自定义Java类:
public class IsaacItem { private String name; private int id; private String iconUrl; private String quote; private String description; private int quality; // 构造方法 public IsaacItem(String name, int id, String iconUrl, String quote, String description, int quality) { this.name = name; this.id = id; this.iconUrl = iconUrl; this.quote = quote; this.description = description; this.quality = quality; } // Getter & Setter 自行补充 }
将items.json放入Android项目的assets目录,用Gson解析成列表:
// 读取assets中的JSON文件 private String loadJsonFromAssets(Context context) { try { InputStream is = context.getAssets().open("items.json"); int size = is.available(); byte[] buffer = new byte[size]; is.read(buffer); is.close(); return new String(buffer, StandardCharsets.UTF_8); } catch (IOException e) { e.printStackTrace(); return null; } } // 生成物品列表 List<IsaacItem> itemList = new Gson().fromJson( loadJsonFromAssets(context), new TypeToken<List<IsaacItem>>() {}.getType() );
方案二:手动整理+硬编码(适合小批量测试)
如果爬虫操作有门槛,直接复制网页表格内容,手动生成列表:
List<IsaacItem> itemList = new ArrayList<>(); itemList.add(new IsaacItem( "Sad Onion", 1, "https://static.wikia.nocookie.net/bindingofisaacre_gamepedia/images/7/7e/Sad_Onion.png", "\"I can cry more tears!\"", "Increases tear rate by 33%.", 1 )); // 逐个添加其他物品
方案三:调用Fandom API获取结构化数据
目标Wiki属于Fandom平台,可直接调用官方API获取结构化物品数据,无需解析HTML表格:
- 调用API接口获取物品列表的JSON响应
- 直接将响应字段映射到你的Java类,稳定性比解析HTML更高
额外提示
- 物品图标无需提前转成位图/文件,可直接用Glide/Picasso加载链接:
Glide.with(context) .load(item.getIconUrl()) .into(itemIconImageView);
- 爬取网页前需遵守网站
robots.txt规则,Fandom允许非商业用途的爬虫请求
内容的提问来源于stack exchange,提问作者miguelrr_11
相关产品推荐
相关产品推荐

