WikiWise is a semantic search engine that enhances information retrieval from Wikipedia movie plot data using Natural Language Processing (NLP). It enables users to search for movie articles intelligently using multiple embedding models and retrieval strategies.
- Semantic Search: Uses multiple embedding models (BERT, BGE, Snowflake, SBERT) to retrieve the most relevant movie articles.
- FAISS Indexing: Supports fast approximate nearest neighbor search using FAISS for efficient retrieval at scale.
- Fine-Tuned Model: Includes a fine-tuned SBERT model for improved domain-specific search accuracy.
- User-Friendly Interface: Simple UI built with Vite, React, TypeScript, and Tailwind CSS for seamless interaction.
- Training and Testing Module: Model training and evaluation handled in a Jupyter Notebook (
Training_and_Testing.ipynb).
/WikiWise-Semantic-Search-Engine
│── data/
│ └── wikipedia_movies_plots.csv
│── frontend/
│ └── (Frontend code using Vite + React + TypeScript + Tailwind CSS)
│── saved_models/
│ └── (Pre-trained model weights and tokenizers)
│── trained_models/
│ └── (Fine-tuned model weights)
│── backend.py
│── Training_and_Testing.ipynb
│── LICENSE
- Frontend: React, TypeScript, Tailwind CSS (Vite for project setup)
- Backend: FastAPI (REST API for handling search queries and model inference)
- NLP Libraries: Hugging Face Transformers, Sentence-Transformers, NLTK
- Search: FAISS (Facebook AI Similarity Search), Cosine Similarity
- ML Frameworks: PyTorch
- Dataset: Wikipedia movie plots CSV
- Jupyter Notebook: Training and testing of models
- Python 3.8+
- Node.js (for frontend development)
-
Clone the Repository:
git clone https://github.com/siftullah/WikiWise-Semantic-Search-Engine.git cd WikiWise-Semantic-Search-Engine -
Backend Setup:
pip install fastapi uvicorn pandas nltk numpy torch scikit-learn transformers sentence-transformers faiss-cpu uvicorn backend:app
The FastAPI server will start at
http://localhost:8000/ -
Frontend Setup:
cd frontend npm install npm run devThe frontend will be accessible at
http://localhost:5173/ -
Training and Testing (Optional):
- Open
Training_and_Testing.ipynbin Jupyter Notebook. - Run the notebook to train and test NLP models.
- Open
- Search Articles: Enter a keyword or phrase in the search bar.
- Model Selection: Choose from multiple embedding models (BERT, BGE, Snowflake, SBERT, SBERT with FAISS, or Fine-Tuned SBERT).
- Results: View ranked movie articles with titles, URLs, and plot text.
Contributions are welcome! To contribute:
- Fork the repository.
- Create a new branch (
feature-branch-name). - Make your changes and commit them.
- Submit a pull request.
This project is licensed under the MIT License.
For any issues or suggestions, feel free to open an issue on GitHub.