Data processing for and with foundation models
A natural language interface for computers
Hub of ready-to-use datasets for ML models
Build AI-powered semantic search applications
Extract schema, statistics and entities from datasets
Training data (data labeling, annotation, workflow) for all data types
ExtractThinker is a Document Intelligence library for LLMs
Data and tools for generating and inspecting OLMo pre-training data
A curated list of data mining papers about fraud detection
DataDreamer: Prompt. Generate Synthetic Data. Train & Align Models
Industrial-strength Natural Language Processing (NLP)
Easy-to-use and powerful NLP library with Awesome model zoo
Superlinked is a Python framework for AI Engineers
Fast and customizable framework for automatic ML model creation
Haystack is an open source NLP framework to interact with your data
Efficient few-shot learning with Sentence Transformers
The Classical Language Toolkit
Toolkit for conversational AI
A Repo For Document AI
Public opinion analysis system
Easy-to-use and high-performance NLP and LLM framework
Dealing with all unstructured data, such as reverse image search
Stanford NLP Python library for many human languages
The most accurate natural language detection library for Python
Data loaders and abstractions for text and NLP