Bert-as-service: Doubt: Can I use this service to obtain docvecs/Paragraph vector of an entire article

Created on 8 Feb 2019  路  4Comments  路  Source: hanxiao/bert-as-service

Hi all,
I am trying to obtain fixed length doc vectors/ Paragraph vectors with this implementation. As mentioned in docs I can increase max_seq_len from 25 to the desired length and pass my article as input. I want to know if this approach is right or is there a downside to it. Also, is there another better approach to obtain docvecs using Bert Model?
Currently, we use gensim library to obtain docvecs for an article. Another approach could be to use word vecs obtained from bert model, one hot encode paragraph ids and obtain its vector (similar to gensim paragraph vec implementation).
What do you guys suggest?

Prerequisites

Please fill in by replacing [ ] with [x].

System information

Some of this information can be collected via this script.

  • OS Platform and Distribution (e.g., Linux Ubuntu 16.04):
  • TensorFlow installed from (source or binary):
  • TensorFlow version:
  • Python version:
  • bert-as-service version:
  • GPU model and memory:
  • CPU model and memory:

Description

Please replace YOUR_SERVER_ARGS and YOUR_CLIENT_ARGS accordingly. You can also write your own description for reproducing the issue.

I'm using this command to start the server:

bert-serving-start YOUR_SERVER_ARGS

and calling the server via:

bc = BertClient(YOUR_CLIENT_ARGS)
bc.encode()

Then this issue shows up:

...

Most helpful comment

I don't think this is the expected use of BERT. BERT is a network trained at the sentence embedding level, thus the representation of more than one sentence should be pretty inaccurate and the computation needed beyond 512 tokens would be huge (remember, the computation isn't linear to the number of tokens).

There's many strategies for you to try if you want a more accurate paragraph representation, for example:

  • do element wise average / max over the sequence of sentence embeddings that compose your paragraphs (no additional training), thus your resulting paragraph embedding will be of the same number of dims than each sentence embedding.
  • use a paragraph embedding from weighted sentences if you have any extra information about how each sentence matters to your overall paragraph representation (no additional training)
  • use a RNN / CNN downstream layer to get a paragraph embedding, however this requires training on a target label (regression / classification task).

All 4 comments

Hi, Currently what I have decided to do is set max_seq_length to 1500 and pass multiple sentences as part of a single sequence separated by ||| but I am encountering
ValueError: The seq length (1500) cannot be greater thanmax_position_embeddings(512).

This is due to the fact that, in https://github.com/google-research/bert/blob/master/modeling.py#L43
the default value of max_position_embeddings is set to 512. Can you open an option to set this from your library as a cmd argument? I can work on this if you want.

I don't think this is the expected use of BERT. BERT is a network trained at the sentence embedding level, thus the representation of more than one sentence should be pretty inaccurate and the computation needed beyond 512 tokens would be huge (remember, the computation isn't linear to the number of tokens).

There's many strategies for you to try if you want a more accurate paragraph representation, for example:

  • do element wise average / max over the sequence of sentence embeddings that compose your paragraphs (no additional training), thus your resulting paragraph embedding will be of the same number of dims than each sentence embedding.
  • use a paragraph embedding from weighted sentences if you have any extra information about how each sentence matters to your overall paragraph representation (no additional training)
  • use a RNN / CNN downstream layer to get a paragraph embedding, however this requires training on a target label (regression / classification task).

length restriction on the server side can now be waived in 1.8.2. This issue is fixed in #236 and the new feature is available since 1.8.2. Please do

pip install bert-serving-client bert-serving-server -U

for the update.

You can now set max_seq_len=NONE when starting a server. In this case, the max_seq_len will be determined by the longest sequence in a batch (or mini-batch if parallelization is activated). That means you can send any sequence shorter than max_position_embeddings (usually 512) defined in bert.json.

You may also want to check the new argument -fixed_embed_length by bert-serving-start --help especially if you intend to use it as ELMo-like embedding.

I don't think this is the expected use of BERT. BERT is a network trained at the sentence embedding level, thus the representation of more than one sentence should be pretty inaccurate and the computation needed beyond 512 tokens would be huge (remember, the computation isn't linear to the number of tokens).

There's many strategies for you to try if you want a more accurate paragraph representation, for example:

  • do element wise average / max over the sequence of sentence embeddings that compose your paragraphs (no additional training), thus your resulting paragraph embedding will be of the same number of dims than each sentence embedding.
  • use a paragraph embedding from weighted sentences if you have any extra information about how each sentence matters to your overall paragraph representation (no additional training)
  • use a RNN / CNN downstream layer to get a paragraph embedding, however this requires training on a target label (regression / classification task).

@ironflood The averaging technique sounds interesting! Could you please point me to the results if this has been tried by anyone?

Was this page helpful?
0 / 5 - 0 ratings

Related issues

lu161513 picture lu161513  路  4Comments

calliwen picture calliwen  路  7Comments

crapthings picture crapthings  路  3Comments

k-oul picture k-oul  路  7Comments

loretoparisi picture loretoparisi  路  6Comments