NLP-Question-Answering

Extracting Data Manually (files in qa_data)

This directory contains question and answer pairs for our testing and training data sets. Each file follows the following standard:

The name of the file is .txt where appears in name_indexer.json (e.g. jon_snow.txt)
the first line is a question
The second line is an answer
The third line is the position of the answer in the context paragraph given by characters.json (e.g. characters[65][paragraphs][0] is the first context paragraph for Jon Snow)
All nonempty lines up until the first empty line reference the first context paragraph. After a new line, all nonempty lines up until the second empty line reference the second context, and so on.
Last line must be an empty line

The format_data.py script will take the data in qa_data and format the data so that the BERT script from huggingface will accept it. This format is how SQuAD formats its data.

Name		Name	Last commit message	Last commit date
Latest commit History 14 Commits
FinScraper		FinScraper
__pycache__		__pycache__
bert_on_colab		bert_on_colab
qa_data		qa_data
scripts		scripts
.gitignore		.gitignore
README.md		README.md
requirements.txt		requirements.txt

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

NLP-Question-Answering

Extracting Data Manually (files in qa_data)

About

Releases

Packages

Contributors 2

Languages

zapper59/NLP-Question-Answering

Folders and files

Latest commit

History

Repository files navigation

NLP-Question-Answering

Extracting Data Manually (files in qa_data)

About

Resources

Stars

Watchers

Forks

Releases

Packages 0

Contributors 2

Languages

Packages