Datasets: NonMatchingChecksumError when loading pubmed dataset

Created on 16 Jun 2020  路  1Comment  路  Source: huggingface/datasets

I get this error when i run nlp.load_dataset('scientific_papers', 'pubmed', split = 'train[:50%]').
The error is:

---------------------------------------------------------------------------
NonMatchingChecksumError                  Traceback (most recent call last)
<ipython-input-2-7742dea167d0> in <module>()
----> 1 df = nlp.load_dataset('scientific_papers', 'pubmed', split = 'train[:50%]')
      2 df = pd.DataFrame(df)
      3 gc.collect()

3 frames
/usr/local/lib/python3.6/dist-packages/nlp/load.py in load_dataset(path, name, version, data_dir, data_files, split, cache_dir, download_config, download_mode, ignore_verifications, save_infos, **config_kwargs)
    518         download_mode=download_mode,
    519         ignore_verifications=ignore_verifications,
--> 520         save_infos=save_infos,
    521     )
    522 

/usr/local/lib/python3.6/dist-packages/nlp/builder.py in download_and_prepare(self, download_config, download_mode, ignore_verifications, save_infos, try_from_hf_gcs, dl_manager, **download_and_prepare_kwargs)
    431                 verify_infos = not save_infos and not ignore_verifications
    432                 self._download_and_prepare(
--> 433                     dl_manager=dl_manager, verify_infos=verify_infos, **download_and_prepare_kwargs
    434                 )
    435                 # Sync info

/usr/local/lib/python3.6/dist-packages/nlp/builder.py in _download_and_prepare(self, dl_manager, verify_infos, **prepare_split_kwargs)
    468         # Checksums verification
    469         if verify_infos:
--> 470             verify_checksums(self.info.download_checksums, dl_manager.get_recorded_sizes_checksums())
    471         for split_generator in split_generators:
    472             if str(split_generator.split_info.name).lower() == "all":

/usr/local/lib/python3.6/dist-packages/nlp/utils/info_utils.py in verify_checksums(expected_checksums, recorded_checksums)
     34     bad_urls = [url for url in expected_checksums if expected_checksums[url] != recorded_checksums[url]]
     35     if len(bad_urls) > 0:
---> 36         raise NonMatchingChecksumError(str(bad_urls))
     37     logger.info("All the checksums matched successfully.")
     38 

NonMatchingChecksumError: ['https://drive.google.com/uc?id=1b3rmCSIoh6VhD4HKWjI4HOW-cSwcwbeC&export=download', 'https://drive.google.com/uc?id=1lvsqvsFi3W-pE1SqNZI0s8NR9rC1tsja&export=download']

I'm currently working on google colab.

That is quite strange because yesterday it was fine.

dataset bug

Most helpful comment

For some reason the files are not available for unauthenticated users right now (like the download service of this package). Instead of downloading the right files, it downloads the html of the error.
According to the error it should be back again in 24h.

image

>All comments

For some reason the files are not available for unauthenticated users right now (like the download service of this package). Instead of downloading the right files, it downloads the html of the error.
According to the error it should be back again in 24h.

image

Was this page helpful?
0 / 5 - 0 ratings

Related issues

astariul picture astariul  路  6Comments

Nouman97 picture Nouman97  路  4Comments

sshleifer picture sshleifer  路  4Comments

thomwolf picture thomwolf  路  4Comments

astariul picture astariul  路  3Comments