Join/Login
Business Software
Open Source Software
For Vendors
Blog
About
More

For Vendors Help Create Join Login

Business Software

Open Source Software

SourceForge Podcast

Resources

Articles
Case Studies
Blog

Menu

Help
Create
Join
Login

Home
Open Source Software
Search Results

Search Results for "web crawler" - Page 4

x

Sort By:

Relevance

OS

Linux 178
Windows 172
Mac 150
More...
BSD 112
ChromeOS 95
Desktop Operating Systems 3
Server Operating Systems 3
Mobile Operating Systems 1

Category

Internet 170
Software Development 22
System 19
Security 16
Communications 9
Scientific/Engineering 8
Database 7
Business 6
Formats and Protocols 5
Artificial Intelligence 4
Desktop Environment 2
Education 2
Games 1
Mobile 1
Multimedia 1
Social sciences 1

License

OSI-Approved Open Source 157
Other License 6
Public Domain 2
Creative Commons Attribution License 1

Translations

English 47
German 7
French 5
Russian 5
More...
Chinese (Simplified) 2
Italian 2
Spanish 2
Brazilian Portuguese 1
Dutch 1
Esperanto 1
Hindi 1
Panjabi 1
Polish 1
Portuguese 1

Programming Language

Python 50
Java 47
PHP 35
JavaScript 24
More...
C++ 16
Perl 12
C# 10
C 9
Go 9
TypeScript 4
Unix Shell 4
ASP 2
PL/SQL 2
PowerShell 2
Visual Basic .NET 2
JSP 1
Kotlin 1
Ruby 1
Rust 1
Scala 1
Visual Basic 1

Status

Production/Stable 40
Beta 25
Alpha 20
Pre-Alpha 18
More...
Planning 10
Mature 1
Inactive 1

Showing 221 open source projects for "web crawler"

View related business solutions

Optimize every aspect of hiring with Greenhouse Recruiting
Hire for what’s next.

What’s next for many of us is changing. Your company’s ability to hire great talent is as important as ever – so you’ll be ready for whatever’s ahead. Whether you need to scale your team quickly or improve your hiring process, Greenhouse gives you the right technology, know-how and support to take on what’s next.

Learn More
SalesTarget.ai | AI-Powered Lead Generation, Email Outreach, and CRM
SalesTarget.ai streamlines your sales process, providing everything you need to find high- quality leads, automate outreach, and close deals faster

SalesTarget is ideal for B2B sales teams, startup founders, and marketing professionals looking to streamline lead generation and outreach. It also benefits growing SaaS companies and agencies aiming to scale their outbound efforts efficiently.

Learn More
1

pxer

Pixiv crawler userscript for downloading artwork and galleries easily

Pxer is an open source tool designed to help users collect and download artwork from the Pixiv illustration platform. It is implemented primarily in client-side JavaScript and runs directly in the browser through a userscript environment, allowing it to integrate seamlessly with Pixiv pages. Pxer provides functionality to crawl and gather images, artwork metadata, and other related content from supported Pixiv pages. It is designed to be accessible even for users who are not developers,...

Downloads: 3 This Week

Last Update: 2026-03-11
See Project
2

mzitu

Python crawler that downloads image galleries and analyzes titles

mzitu is a Python-based web crawling project designed to automatically download and organize image galleries from a specific photography site. It demonstrates how to build a scraper that navigates gallery pages, retrieves image links, and saves the images locally in a structured directory layout. It focuses on automating the collection of large sets of images by programmatically parsing page content and iterating through gallery entries. mzitu also includes a simple analysis script that...

Downloads: 1 This Week

Last Update: 1 day ago
See Project
3

WeChatSogou

Python library to crawl and retrieve data from WeChat accounts

WechatSogou is an open source Python library designed to retrieve data from WeChat official accounts by using the Sogou WeChat search service as its data source. It provides developers with a programmatic way to search for public accounts and collect article information without manually browsing the search interface. It functions as a crawler interface that sends requests to the search engine, retrieves results, and converts the returned pages into structured data that can be used in...

Downloads: 6 This Week

Last Update: 2026-03-10
See Project
4

OpenSearchServer Search Engine

An open source search engine with RESTFul API and crawlers

OpenSearchServer is a powerful, enterprise-class, search engine program. Using the web user interface, the crawlers (web, file, database, etc.) and the client libraries (REST/API , Ruby, Rails, Node.js, PHP, Perl) you will be able to integrate quickly and easily advanced full-text search capabilities in your application: Full-text with basic semantic, join queries, boolean queries, facet and filter, document (PDF, Office, etc.) indexation, web scrapping,etc. OpenSearchServer runs on...

31 Reviews

Downloads: 12 This Week

Last Update: 2018-08-26
See Project
Jscrambler: Pioneering Client-Side Protection Platform
Jscrambler offers an exclusive blend of cutting-edge first-party JavaScript obfuscation and state-of-the-art third-party tag protection.

Jscrambler is the leader in Client-Side Protection and Compliance. We were the first to merge advanced polymorphic JavaScript obfuscation with fine-grained third-party tag protection in a unified Client-Side Protection and Compliance Platform. Our integrated solution ensures a robust defense against current and emerging client-side cyber threats, data leaks, and IP theft, empowering software development and digital teams to innovate securely. With Jscrambler, businesses adopt a unified, future-proof client-side security policy all while achieving compliance with emerging security standards including PCI DSS v4.0. Trusted by digital leaders worldwide, Jscrambler gives businesses the freedom to innovate securely.

Learn More
5

crawler4j

Open source web crawler for Java

crawler4j is an open source web crawler for Java which provides a simple interface for crawling the Web. Using it, you can setup a multi-threaded web crawler in few minutes. You need to create a crawler class that extends WebCrawler. This class decides which URLs should be crawled and handles the downloaded page. shouldVisit function decides whether the given URL should be crawled or not.

Downloads: 0 This Week

Last Update: 2022-01-12
See Project
6

pyspider

A powerful Spider(Web Crawler) system in Python

pyspider is a powerful Spider(Web Crawler) system in Python. Components are connected by message queue. Every component, including message queue, is running in their own process/thread, and replaceable. That means, when process is slow, you can have many instances of processor and make full use of multiple CPUs, or deploy to multiple machines. This architecture makes pyspider really fast. benchmarking.

Downloads: 0 This Week

Last Update: 2021-03-31
See Project
7

haipproxy

Distributed proxy IP pool for web crawlers using Scrapy and Redis

...HAipproxy aims to maintain a high availability proxy pool with low latency so that scraping frameworks can rotate proxies efficiently and avoid blocking during large-scale data collection. Its architecture supports distributed deployment, allowing multiple crawler workers and validators to run across different machines.

Downloads: 0 This Week

Last Update: 2026-03-10
See Project
8

gain

Asyncio-based Python framework for building fast web crawling spiders

Gain is a Python web crawling framework designed to simplify the process of building efficient and scalable web scrapers. It is built on top of asynchronous technologies such as asyncio, aiohttp, and uvloop to support high-performance crawling with concurrent network requests. It provides a structured framework for creating spiders that can navigate websites, extract structured data, and process the collected results. Developers define crawlers using components such as spiders, parsers, and...

Downloads: 1 This Week

Last Update: 1 day ago
See Project
9

Toapi

Convert websites into structured APIs automatically with Python tool

Toapi is a Python library designed to transform ordinary websites into usable API services. Instead of building a traditional web crawler that collects and stores data before exposing it through an API, Toapi simplifies the process by allowing developers to define data structures that automatically generate an API layer from existing web pages. It works by parsing HTML content from a source site and mapping selected elements into structured data that can be returned as JSON through API endpoints. ...

Downloads: 0 This Week

Last Update: 1 day ago
See Project
Powering the next decade of business messaging | Twilio MessagingX
For organizations interested programmable APIs built on a scalable business messaging platform

Build unique experiences across SMS, MMS, Facebook Messenger, and WhatsApp – with our unified messaging APIs.

Learn More
10

Gecco

Lightweight Java web crawler framework with jQuery-style extraction

Gecco is a lightweight web crawler framework written in Java that simplifies the process of building web scraping applications. It is designed to make crawler development straightforward by allowing developers to extract page elements using jQuery-style selectors rather than complex parsing logic. It integrates several well-known Java libraries and frameworks, including tools for HTTP requests, HTML parsing, JSON processing, and application development. ...

Downloads: 3 This Week

Last Update: 2026-03-10
See Project
11

diskover

File system crawler and disk space usage software

diskover is a file system crawler and disk space usage software that uses Elasticsearch to index your file metadata. diskover crawls and indexes your files on a local computer or remote storage server over network mounts. diskover helps manage your storage by identifying old and unused files and give better insights into data change "hotfiles", file duplication "dupes" and wasted space. It is designed to help deal with managing large amounts of data growth and provide detailed storage...

Downloads: 0 This Week

Last Update: 2020-05-16
See Project
12

Perl Web Scraping Project

Perl Web Scraping Project

Web scraping (web harvesting or web data extraction) is data scraping used for extracting data from websites.[1] Web scraping software may access the World Wide Web directly using the Hypertext Transfer Protocol, or through a web browser. While web scraping can be done manually by a software user, the term typically refers to automated processes implemented using a bot or web crawler.

Downloads: 0 This Week

Last Update: 2017-10-12
See Project
13

lightcrawler

Website crawler that audits site pages automatically with Lighthouse

...It works by starting from a given URL and recursively exploring linked pages to collect a set of pages that should be analyzed. Each discovered page is then evaluated using Lighthouse, which performs checks related to performance, accessibility, and web development best practices. This allows developers to audit multiple pages of a site automatically instead of manually running Lighthouse on each individual page. Lightcrawler supports configuration through a JSON configuration file, enabling users to customize how the crawler operates and which Lighthouse audits should be executed. ...

Downloads: 16 This Week

Last Update: 5 days ago
See Project
14

Catberry

Catberry is an isomorphic framework

Catberry is an isomorphic framework for building universal front-end apps using components, Flux architecture and progressive rendering. Catberry builds a bundle for running the application in a browser as a Single Page Application. Cat-Components – similar to web-components but organized as directories, can be rendered on the server and published/installed as NPM packages. The entire architecture of the framework is built using the Service Locator pattern, which helps to manage module dependencies and create plugins, and Flux, for the data layer. Search crawler receives a full page from the server. ...

Downloads: 0 This Week

Last Update: 2022-12-01
See Project
15

phoneutria

A Java Web crawler: multi-threaded, scalable, with high performance, extensible and polite. It can be used to crawl and index any web or enterprise domain and is configurable through a XML configuration file.

Downloads: 0 This Week

Last Update: 2017-05-22
See Project
16

OpenWebSpider

OpenWebSpider is an Open Source multi-threaded Web Spider (robot, crawler) and search engine with a lot of interesting features!

4 Reviews

Downloads: 5 This Week

Last Update: 2017-03-12
See Project
17

Solr Web Crawler

Downloads: 0 This Week

Last Update: 2016-10-26
See Project
18

sourcegreed

a java-based crawler

a java-based crawler

Downloads: 0 This Week

Last Update: 2016-07-27
See Project
19

WebCrawler

get web page. include html、css and js files

This tool is for the people who want to learn from a web site or web page,especially Web Developer.It can help get a web page's source code.Input the web page's address and press start button and this tool will find the page and according the page's quote,download all files that used in the page ,include css file and javascript files. The html file's name will be 'index.html' and other file's will use it's source name. Note:only support windows platform and http protocol.

Downloads: 2 This Week

Last Update: 2016-04-16
See Project
20

WebCollector

WebCollector is an open source web crawler framework based on Java.

WebCollector is an open source web crawler framework based on Java.It provides some simple interfaces for crawling the Web,you can setup a multi-threaded web crawler in less than 5 minutes. Github: https://github.com/CrawlScript/WebCollector Demo: https://github.com/CrawlScript/WebCollector/blob/master/YahooCrawler.java

Downloads: 0 This Week

Last Update: 2015-06-04
See Project
21

htmlparser

Products of the project: Java HTMLParser - VietSpider Web Data Extractor - Extractor VietSpider News. Click on "Show project details" to see more feature about each product.

Downloads: 0 This Week

Last Update: 2015-06-24
See Project
22

Node Crawler

Web Crawler/Spider for NodeJS + server-side jQuery

Most powerful, popular and production crawling/scraping package for Node, happy hacking.

Downloads: 2 This Week

Last Update: 2023-09-20
See Project
23

crawly

Simple website crawler

soon.

Downloads: 0 This Week

Last Update: 2014-12-02
See Project
24

Addons for IOSEC - DoS HTTP Security

IOSec Addons are enhancements for web security and crawler detection

IOSEC PHP HTTP FLOOD PROTECTION ADDONS IOSEC is a php component that allows you to simply block unwanted access to your webpage. if a bad crawler uses to much of your servers resources iosec can block that. IOSec Enhanced...

2 Reviews

Downloads: 5 This Week

Last Update: 2023-04-26
See Project
25

GHZ Tools

GHZ Tools v0.6 Build 9645 Release Data (02/09/2014) 7zPass: MHg2NzY4N0E3NDZGNkY2QzczMzAzNj== (base64/hex) Properties: 1)- Brute Forcer: WordPress Joomla 4images osCommerce Drupal, Razor Ftp cPanel Whmcs DirectAdmin Authentication Bypass SSH Authentication vBulletin Kleeja OpenCart WordPress Xmlrpc 2)- Remote Exploits: JCE Webdav 3)- SQL Injector: Auto SQL Injection 4)- Hash Cracker: MD2 MD4 MD5 SHA1 MD5(MD5(PASS)) SHA1(SHA1(PASS)) 5)- URL Fuzzer: URL Fuzzer 6)- Web Scanner: RFI/LFI URL Scanner Web Extractor Open Port Scanner URL Crawler SQLi Scanner

Downloads: 0 This Week

Last Update: 2014-09-02
See Project

Previous
1
2
3
You're on page 4
5
6
7
8
9
Next

Related Searches

torrent search

web crawler

lg bypass tool

web scraping

news crawler

web spider

sql injector

torrent

gitst web crawler

dark web scraper

Related Categories

Internet

Software Development

System

Security

Communications

SourceForge

Create a Project
Open Source Software
Business Software
Top Downloaded Projects

Company

About
Team
SourceForge Headquarters
1320 Columbia Street Suite 310
San Diego, CA 92101
+1 (858) 422-6466

Resources

Support
Site Documentation
Site Status
SourceForge Reviews

© 2026 Slashdot Media. All Rights Reserved.

Terms Privacy Privacy Choices Advertise