# Collect html text content HELP!

**URL:** <https://community.thunkable.com/t/collect-html-text-content-help/1173096>\
**Category:** Questions about Thunkable X\
**Created:** [March 17, 2021, 9:26pm UTC](https://community.thunkable.com/t/collect-html-text-content-help/1173096 "2021-03-17T21:26:38Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Fernando\_Matos](https://sea1.discourse-cdn.com/flex015/user_avatar/community.thunkable.com/fernando_matos/32/30641_2.png) [@Fernando\_Matos](https://community.thunkable.com/u/Fernando_Matos)\
**Post date:** [March 17, 2021, 9:26pm UTC](https://community.thunkable.com/t/collect-html-text-content-help/1173096/1 "2021-03-17T21:26:38Z")

</div>

Hello thunkers, I have a problem in my hands to solve a function in my new application, and I am in need of help. i am using webApi to receive the html code from a web site. the html code and then shown in a label. now the next one i needed to create a list of some information that appears inside the html code. all codes I want start with “1” and are 20 characters long. how can i collect this data from the html code and create a list? thanks

---

<div class="post-metadata">

**Author:** ![tatiang](https://sea1.discourse-cdn.com/flex015/user_avatar/community.thunkable.com/tatiang/32/55482_2.png) [@tatiang](https://community.thunkable.com/u/tatiang)\
**Post date:** [March 17, 2021, 10:54pm UTC](https://community.thunkable.com/t/collect-html-text-content-help/1173096/2 "2021-03-17T22:54:16Z")

</div>

What does the html look like? It sounds like you will need to parse it but I can’t explain how to do that without seeing it.

---

<div class="post-metadata">

**Author:** ![Fernando\_Matos](https://sea1.discourse-cdn.com/flex015/user_avatar/community.thunkable.com/fernando_matos/32/30641_2.png) [@Fernando\_Matos](https://community.thunkable.com/u/Fernando_Matos)\
**Post date:** [March 20, 2021, 12:23am UTC](https://community.thunkable.com/t/collect-html-text-content-help/1173096/3 "2021-03-20T00:23:44Z")

</div>

Hello. the html contains private information that I don’t want to disclose. what I want is exactly what is shown in the topic below. but in the example it is only for one line of text, and what I intend to do is the same for 128 lines of which everyone starts with 1. Thanks

> [@Parsing HTML in Thunkable ✕ Cross Platform](https://community.thunkable.com/t/parsing-html-in-thunkable-cross-platform/115177):
>
> Hello! Help me please, is it possible to parse HTML pages in Thunkable ✕ Cross Platform? I need to get some text from the HTML page that contains: {«status»: 200, «data»: «486F6D6553616D6F676F6C312C3737372E302C3737372E302C302C»} I know that the text starts at «status»: 200, «data»: « and ends at 302C»}, can I get this part of the text: 486F6D6553616D6F676F6C312C3737372E302C3737372E302C302C ? Thanks in advance!

---

<div class="post-metadata">

**Author:** ![Fernando\_Matos](https://sea1.discourse-cdn.com/flex015/user_avatar/community.thunkable.com/fernando_matos/32/30641_2.png) [@Fernando\_Matos](https://community.thunkable.com/u/Fernando_Matos)\
**Post date:** [March 23, 2021, 12:19am UTC](https://community.thunkable.com/t/collect-html-text-content-help/1173096/4 "2021-03-23T00:19:43Z")

</div>

Some HELP??

---

<div class="post-metadata">

**Author:** ![japa6225a](https://sea1.discourse-cdn.com/flex015/user_avatar/community.thunkable.com/japa6225a/32/95982_2.png) [@japa6225a](https://community.thunkable.com/u/japa6225a)\
**Post date:** [March 23, 2021, 12:36am UTC](https://community.thunkable.com/t/collect-html-text-content-help/1173096/5 "2021-03-23T00:36:49Z")

</div>

That code is JSON. use the ![image](https://us1.discourse-cdn.com/flex015/uploads/thunkable/original/3X/6/4/64083956824222bb3877157c8e4c0e3295670cd9.png)  
Note: ![image](https://us1.discourse-cdn.com/flex015/uploads/thunkable/original/3X/4/5/451afaa283f7bf9dc089cfeeb44c653a677b0b72.png)

---

<div class="post-metadata">

**Author:** ![tatiang](https://sea1.discourse-cdn.com/flex015/user_avatar/community.thunkable.com/tatiang/32/55482_2.png) [@tatiang](https://community.thunkable.com/u/tatiang)\
**Post date:** [March 23, 2021, 2:35am UTC](https://community.thunkable.com/t/collect-html-text-content-help/1173096/6 "2021-03-23T02:35:09Z")

</div>

I’m guessing you’ll want to do something like this but it’s impossible to know without seeing at least a sample (with private data modified):

> [@Parsing HTML in Thunkable ✕ Cross Platform](https://community.thunkable.com/t/parsing-html-in-thunkable-cross-platform/115177/2):
>
> Hello!

---

<div class="post-metadata">

**Author:** ![Fernando\_Matos](https://sea1.discourse-cdn.com/flex015/user_avatar/community.thunkable.com/fernando_matos/32/30641_2.png) [@Fernando\_Matos](https://community.thunkable.com/u/Fernando_Matos)\
**Post date:** [March 25, 2021, 12:12am UTC](https://community.thunkable.com/t/collect-html-text-content-help/1173096/7 "2021-03-25T00:12:19Z")

</div>

well, this is the beginning of the html i intend to create a list of the selected items in green. the code contains dozens of them. and I know that they all start with 1 and have the same number of characters

 ![image](https://us1.discourse-cdn.com/flex015/uploads/thunkable/original/3X/0/2/02d277adc4f88c0570edcd07a43ce47183261a5a.jpeg)

---

<div class="post-metadata">

**Author:** ![catsarisky](https://sea1.discourse-cdn.com/flex015/user_avatar/community.thunkable.com/catsarisky/32/100661_2.png) [@catsarisky](https://community.thunkable.com/u/catsarisky)\
**Post date:** [March 25, 2021, 2:19am UTC](https://community.thunkable.com/t/collect-html-text-content-help/1173096/8 "2021-03-25T02:19:42Z")

</div>

I think the answer is probably in that link tatiang posted, but you’ll need to do some work to use it.  
Is the rest of the page pretty constant? Content could change, but formatting is going to be rigidly the same? (If a computer is producing that output, that’s probably true. If a human is whacking at HTML to update the page, probably not.)  
Look for what delimits the content you want (is it one per chunk that starts with span\_id=“whatever’s under the red” ?) and use list splitting to split your content into chunks. Then parse each one, either by length (if the content you want is rigidly the same length) or by finding more delimiters for splitting it up.

---

<div class="post-metadata">

**Author:** ![catsarisky](https://sea1.discourse-cdn.com/flex015/user_avatar/community.thunkable.com/catsarisky/32/100661_2.png) [@catsarisky](https://community.thunkable.com/u/catsarisky)\
**Post date:** [March 25, 2021, 3:02am UTC](https://community.thunkable.com/t/collect-html-text-content-help/1173096/9 "2021-03-25T03:02:05Z")

</div>

![image](https://us1.discourse-cdn.com/flex015/uploads/thunkable/original/3X/4/6/46319afe5afa2c5a8e831c4efec78af1a8f2d9e2.png)

The sample code above reads the html from [community.thunkable.com](http://community.thunkable.com) and pulls out a list of categories. You can probably adapt it to get out those strings you want, IF the rest of the webpage is pretty consistently formatted.

Here’s the link, if you’d like to remix: [Thunkable](https://x.thunkable.com/copy/cc96507414c6be6b2ebf81b8afbcad5d)

---

<div class="post-metadata">

**Author:** ![ioannis](https://sea1.discourse-cdn.com/flex015/user_avatar/community.thunkable.com/ioannis/32/146956_2.png) [@ioannis](https://community.thunkable.com/u/ioannis)\
**Post date:** [November 8, 2024, 1:25pm UTC](https://community.thunkable.com/t/collect-html-text-content-help/1173096/10 "2024-11-08T13:25:57Z")

</div>


