discourse

Commit Graph

Author	SHA1	Message	Date
Martin Brennan	b500949ef6	FEATURE: Initial implementation of direct S3 uploads with uppy and stubs (#13787 ) This adds a few different things to allow for direct S3 uploads using uppy. These changes are still not the default. There are hidden `enable_experimental_image_uploader` and `enable_direct_s3_uploads` settings that must be turned on for any of this code to be used, and even if they are turned on only the User Card Background for the user profile actually uses uppy-image-uploader. A new `ExternalUploadStub` model and database table is introduced in this pull request. This is used to keep track of uploads that are uploaded to a temporary location in S3 with the direct to S3 code, and they are eventually deleted a) when the direct upload is completed and b) after a certain time period of not being used. ### Starting a direct S3 upload When an S3 direct upload is initiated with uppy, we first request a presigned PUT URL from the new `generate-presigned-put` endpoint in `UploadsController`. This generates an S3 key in the `temp` folder inside the correct bucket path, along with any metadata from the clientside (e.g. the SHA1 checksum described below). This will also create an `ExternalUploadStub` and store the details of the temp object key and the file being uploaded. Once the clientside has this URL, uppy will upload the file direct to S3 using the presigned URL. Once the upload is complete we go to the next stage. ### Completing a direct S3 upload Once the upload to S3 is done we call the new `complete-external-upload` route with the unique identifier of the `ExternalUploadStub` created earlier. Only the user who made the stub can complete the external upload. One of two paths is followed via the `ExternalUploadManager`. 1. If the object in S3 is too large (currently 100mb defined by `ExternalUploadManager::DOWNLOAD_LIMIT`) we do not download and generate the SHA1 for that file. Instead we create the `Upload` record via `UploadCreator` and simply copy it to its final destination on S3 then delete the initial temp file. Several modifications to `UploadCreator` have been made to accommodate this. 2. If the object in S3 is small enough, we download it. When the temporary S3 file is downloaded, we compare the SHA1 checksum generated by the browser with the actual SHA1 checksum of the file generated by ruby. The browser SHA1 checksum is stored on the object in S3 with metadata, and is generated via the `UppyChecksum` plugin. Keep in mind that some browsers will not generate this due to compatibility or other issues. We then follow the normal `UploadCreator` path with one exception. To cut down on having to re-upload the file again, if there are no changes (such as resizing etc) to the file in `UploadCreator` we follow the same copy + delete temp path that we do for files that are too large. 3. Finally we return the serialized upload record back to the client There are several errors that could happen that are handled by `UploadsController` as well. Also in this PR is some refactoring of `displayErrorForUpload` to handle both uppy and jquery file uploader errors.	2021-07-28 08:42:25 +10:00
Gerhard Schlager	157f10db4c	FEATURE: Use path from existing URL of uploads and optimized images (#13177 ) Discourse shouldn't dynamically calculate the path of uploads and optimized images after a file has been stored on disk or S3. Otherwise it might calculate the wrong path if the SHA1 or extension stored in the database doesn't match the actual file path.	2021-05-27 17:42:25 +02:00
Josh Soref	59097b207f	DEV: Correct typos and spelling mistakes (#12812 ) Over the years we accrued many spelling mistakes in the code base. This PR attempts to fix spelling mistakes and typos in all areas of the code that are extremely safe to change - comments - test descriptions - other low risk areas	2021-05-21 11:43:47 +10:00
David Taylor	13e39d8b9f	PERF: Improve cook_url performance for topic thumbnails (#11609 ) - Only initialize the S3Helper when needed - Skip initializing the S3Helper for S3Store#cdn_url - Allow cook_url to be passed a `local` hint to skip unnecessary checks	2020-12-30 18:13:13 +00:00
Martin Brennan	4193eb0419	FIX: Respect force download when downloading secure media via lightbox (#10769 ) The download link on the lightbox for images was not downloading the image if the upload was marked secure, because the code in the upload controller route was not respecting the dl=1 param for force download. This PR fixes this so the download link works for secure images as well as regular ligthboxed images.	2020-09-29 12:12:03 +10:00
Martin Brennan	31e31ef449	SECURITY: Add content-disposition: attachment for SVG uploads * strip out the href and xlink:href attributes from use element that are _not_ anchors in svgs which can be used for XSS * adding the content-disposition: attachment ensures that uploaded SVGs cannot be opened and executed using the XSS exploit. svgs embedded using an img tag do not suffer from the same exploit	2020-07-09 13:31:48 +10:00
Martin Brennan	8ef782bdbd	FIX: Increase time of DOWNLOAD_URL_EXPIRES_AFTER_SECONDS to 5 minutes (#10160 ) * Change S3Helper::DOWNLOAD_URL_EXPIRES_AFTER_SECONDS to 5 minutes, which controls presigned URL expiry and secure-media route cache time. * This is done because of the composer preview refreshing while typing causes a lot of requests sent to our server because of the short URL expiry. If this ends up being not enough we can always increase the time or explore other avenues (e.g. GitHub has a 7 day validity for secure URLs)	2020-07-03 13:42:36 +10:00
Sam Saffron	689568c216	FIX: invalid urls should not break store.has_been_uploaded? Breaking this method has wide ramification including breaking search indexing.	2020-06-25 15:00:15 +10:00
Martin Brennan	e92909aa77	FIX: Use ActionDispatch::Http::ContentDisposition for uploads content-disposition (#10108 ) See https://meta.discourse.org/t/broken-pipe-error-when-uploading-to-a-s3-clone-a-pdf-with-a-name-containing-e-i-etc/155414 When setting content-disposition for attachment, use the ContentDisposition class to format it. This handles filenames with weird characters and localization (accented characters) correctly.	2020-06-23 17:10:56 +10:00
Guo Xiang Tan	828ceab64b	DEV: Make rubocop happy.	2020-06-17 15:47:05 +08:00
Martin Brennan	e5da2d24e5	FIX: Add attachment content-disposition for all non-image files (#10058 ) This will make it so the original filename is used when downloading all non-image files, bringing S3Store into line with the to_s3 migration and local storage. Video and audio files will still stream correctly in HTML players as well. See https://meta.discourse.org/t/cannot-download-non-image-media-files-original-filenames-lost-when-uploaded-to-s3/152797 for a lot of extra context.	2020-06-17 11:16:37 +10:00
Roman Rizzi	b61a291cf3	FIX: returns false if the upload url is an invalid mailto link (#9877 )	2020-05-26 10:32:48 -03:00
Michael Brown	d9a02d1336	Revert "Revert "Merge branch 'master' of https://github.com/discourse/discourse "" This reverts commit `20780a1eee`. * SECURITY: re-adds accidentally reverted commit: 03d26cd6: ensure embed_url contains valid http(s) uri * when the merge commit `e62a85cf` was reverted, git chose the `2660c2e2` parent to land on instead of the `03d26cd6` parent (which contains security fixes)	2020-05-23 00:56:13 -04:00
Jeff Atwood	20780a1eee	Revert "Merge branch 'master' of https://github.com/discourse/discourse " This reverts commit `e62a85cf6f`, reversing changes made to `2660c2e21d`.	2020-05-22 20:25:56 -07:00
Osama Sayegh	02f44def56	FIX: Don't blow up when trying to parse invalid or non-ASCII URLs (#9838 ) * FIX: Don't blow up when trying to parseinvalid or non-ASCII URLs Follow-up to `72f139191e`	2020-05-20 12:46:27 +03:00
Martin Brennan	72f139191e	FIX: S3 store has_been_uploaded? was not taking into account s3 bucket path (#9810 ) In some cases, between Discourse forums the hostname of a URL could match if they are hosting S3 files on the same bucket but the S3 bucket path might not. So e.g. https://testbucket.somesite.com/testpath/some/file/url.png vs https://testbucket.somesite.com/prodpath/some/file/url.png. So has_been_uploaded? was returning true for the second URL, even though it may have been uploaded on a different Discourse forum. This is a very rare case but must be accounted for, because this impacts UrlHelper.is_local which mistakenly thinks the file has already been downloaded and thus allows the URL to be cooked, where we want to return the full URL to be downloaded using PullHotlinkedImages.	2020-05-20 10:40:38 +10:00
Gerhard Schlager	c6b411f6c1	FIX: Restore to S3 didn't work without env variables The `uplaods:migrate_to_s3` rake task should always use the environment variables, because you usually don't want to break your site's uploads during the migration. But restoring a backup should work with site settings as well as environment variables, otherwise you can't restore uploads to S3 from the web interface.	2020-04-19 20:24:40 +02:00
Martin Brennan	7c32411881	FEATURE: Secure media allowing duplicated uploads with category-level privacy and post-based access rules (#8664 ) ### General Changes and Duplication * We now consider a post `with_secure_media?` if it is in a read-restricted category. * When uploading we now set an upload's secure status straight away. * When uploading if `SiteSetting.secure_media` is enabled, we do not check to see if the upload already exists using the `sha1` digest of the upload. The `sha1` column of the upload is filled with a `SecureRandom.hex(20)` value which is the same length as `Upload::SHA1_LENGTH`. The `original_sha1` column is filled with the _real_ sha1 digest of the file. * Whether an upload `should_be_secure?` is now determined by whether the `access_control_post` is `with_secure_media?` (if there is no access control post then we leave the secure status as is). * When serializing the upload, we now cook the URL if the upload is secure. This is so it shows up correctly in the composer preview, because we set secure status on upload. ### Viewing Secure Media * The secure-media-upload URL will take the post that the upload is attached to into account via `Guardian.can_see?` for access permissions * If there is no `access_control_post` then we just deliver the media. This should be a rare occurrance and shouldn't cause issues as the `access_control_post` is set when `link_post_uploads` is called via `CookedPostProcessor` ### Removed We no longer do any of these because we do not reuse uploads by sha1 if secure media is enabled. * We no longer have a way to prevent cross-posting of a secure upload from a private context to a public context. * We no longer have to set `secure: false` for uploads when uploading for a theme component.	2020-01-16 13:50:27 +10:00
Gerhard Schlager	e474cda321	REFACTOR: Restoring of backups and migration of uploads to S3	2020-01-14 11:41:35 +01:00
Penar Musaraj	102909edb3	FEATURE: Add support for secure media (#7888 ) This PR introduces a new secure media setting. When enabled, it prevent unathorized access to media uploads (files of type image, video and audio). When the `login_required` setting is enabled, then all media uploads will be protected from unauthorized (anonymous) access. When `login_required`is disabled, only media in private messages will be protected from unauthorized access. A few notes: - the `prevent_anons_from_downloading_files` setting no longer applies to audio and video uploads - the `secure_media` setting can only be enabled if S3 uploads are already enabled and configured - upload records have a new column, `secure`, which is a boolean `true/false` of the upload's secure status - when creating a public post with an upload that has already been uploaded and is marked as secure, the post creator will raise an error - when enabling or disabling the setting on a site with existing uploads, the rake task `uploads:ensure_correct_acl` should be used to update all uploads' secure status and their ACL on S3	2019-11-18 11:25:42 +10:00
Penar Musaraj	067696df8f	DEV: Apply Rubocop redundant return style	2019-11-14 15:10:51 -05:00
Daniel Waterworth	55a1394342	DEV: pluck_first Doing .pluck(:column).first is a very common pattern in Discourse and in most cases, a limit cause isn't being added. Instead of adding a limit clause to all these callsites, this commit adds two new methods to ActiveRecord::Relation: pluck_first, equivalent to limit(1).pluck(*columns).first and pluck_first! which, like other finder methods, raises an exception when no record is found	2019-10-21 12:08:20 +01:00
Gerhard Schlager	24877a7b8c	FIX: Correctly encode non-ASCII filenames in HTTP header Backport of fix from Rails 6: `890485cfce`	2019-08-07 19:10:50 +02:00
Rafael dos Santos Silva	606c0ed14d	FIX: S3 uploads were missing a cache-control header (#7902 ) Admins still need to run the rake task to fix the files who where uploaded previously.	2019-08-06 14:55:17 -03:00
Gerhard Schlager	f2dc59d61f	FEATURE: Add hidden setting to include S3 uploads in backups	2019-07-09 14:04:16 +02:00
Penar Musaraj	03805e5a76	FIX: Ensure lightbox image download has correct content disposition in S3 (#7845 )	2019-07-04 11:32:51 -04:00
Penar Musaraj	f00275ded3	FEATURE: Support private attachments when using S3 storage (#7677 ) * Support private uploads in S3 * Use localStore for local avatars * Add job to update private upload ACL on S3 * Test multisite paths * update ACL for private uploads in migrate_to_s3 task	2019-06-06 13:27:24 +10:00
Guo Xiang Tan	a3938f98f8	Revert changes to `FileStore::S3Store#path_for` in `f0620e7118`. There are some places in the code base that assumes the method should return nil.	2019-05-29 18:39:07 +08:00
Guo Xiang Tan	f0620e7118	FEATURE: Support `[description\|attachment](upload://<short-sha>)` in MD take 2. Previous attempt was missing `post_uploads` records.	2019-05-29 09:26:32 +08:00
Penar Musaraj	7c9fb95c15	Temporarily revert "FEATURE: Support `[description\|attachment](upload://<short-sha>)` in MD. (#7603 )" This reverts commit `b1d3c678ca`. We need to make sure post_upload records are correctly stored.	2019-05-28 16:37:01 -04:00
Guo Xiang Tan	b1d3c678ca	FEATURE: Support `[description\|attachment](upload://<short-sha>)` in MD. (#7603 )	2019-05-28 11:18:21 -04:00
Sam Saffron	30990006a9	DEV: enable frozen string literal on all files This reduces chances of errors where consumers of strings mutate inputs and reduces memory usage of the app. Test suite passes now, but there may be some stuff left, so we will run a few sites on a branch prior to merging	2019-05-13 09:31:32 +08:00
Guo Xiang Tan	243fb8d9ad	Fix the build.	2019-03-13 17:39:07 +08:00
Vinoth Kannan	563b953224	DEV: Add 'backfill_etags_' to the method name since it also backfilling the etags	2019-02-19 21:54:35 +05:30
Vinoth Kannan	0472bd4adc	FIX: Remove 'backfill_etags' keyword argument from 'uploads:missing' rake task And etags backfilling code is optimized	2019-02-15 00:34:35 +05:30
Vinoth Kannan	7b5931013a	Update rake task to backfill etags from s3 inventory	2019-02-14 05:18:06 +05:30
Vinoth Kannan	b4f713ca52	FEATURE: Use amazon s3 inventory to manage upload stats (#6867 )	2019-02-01 10:10:48 +05:30
Vinoth Kannan	75dbb98cca	FEATURE: Add S3 etag value to uploads table (#6795 )	2019-01-04 14:16:22 +08:00
Rishabh	cae5ba7356	FIX: Ensure that multisite s3 uploads are tombstoned correctly (#6769 ) * FIX: Ensure that multisite uploads are tombstoned into the correct paths * Move multisite specs to spec/multisite/s3_store_spec.rb	2018-12-19 13:32:32 +08:00
Rishabh	503ae1829f	FIX: All multisite upload paths should start with /uploads/default/.. (#6707 )	2018-12-03 12:04:14 +08:00
Rishabh	05a4f3fb51	FEATURE: Multisite support for S3 image stores (#6689 ) * FEATURE: Multisite support for S3 image stores * Use File.join to concatenate all paths & fix linting on multisite/s3_store_spec.rb	2018-11-29 12:11:48 +08:00
Vinoth Kannan	bcdf5b2f47	DEV: improve missing uploads query and skip checking file size	2018-11-27 02:21:33 +05:30
Vinoth Kannan	4ccf9d28eb	Remove trailing whitespaces	2018-11-27 01:15:29 +05:30
Vinoth Kannan	fd272eee44	FEATURE: Make uploads:missing task compatible with s3 uploads	2018-11-27 00:54:51 +05:30
Guo Xiang Tan	e1b16e445e	Rename `FileHelper.is_image?` -> `FileHelper.is_supported_image?`.	2018-09-12 09:22:28 +08:00
Guo Xiang Tan	8496537590	Add `RECOVER_FROM_S3` to `uploads:list_posts_with_broken_images` rake task.	2018-09-10 15:14:30 +08:00
Sam	5d96809abd	FIX: improve support for subfolder S3 CDN	2018-08-22 12:31:13 +10:00
Sam	f5142861e5	Revert "Revert "FIX: upload URLs from S3 on subfolder installs"" This reverts commit `26c96e97e5`. We have no choice but to run this code	2018-08-22 11:31:33 +10:00
Sam	26c96e97e5	Revert "FIX: upload URLs from S3 on subfolder installs" This reverts commit `357df2ff4f`.	2018-08-22 10:51:40 +10:00
Neil Lalonde	357df2ff4f	FIX: upload URLs from S3 on subfolder installs	2018-08-21 14:58:55 -04:00

1 2

93 Commits